@ccoalm/ccl-skills 0.1.0 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_init_policy_matrix.sh +93 -16
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_probe_result.sh +10 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +249 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate_abort_leak.sh +394 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/cross-cutting-concerns.md +6 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +3 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-capability-composition.md +128 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-command-sandbox.md +35 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-session-persistence.md +54 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-tool-dispatch.md +11 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/model-prompt-evaluation.md +7 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/canary-and-rollout-strategy.md +6 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/external-ui-ux-quality-benchmarks.md +50 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-audit.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/architecture-playbook.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/packaging-runtime-readiness.md +6 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +59 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/rule-consolidation.md +3 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +43 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +16 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +391 -14
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb +74 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +74 -33
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh +11 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_verdict_differential.sh +421 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_round_attribution.sh +576 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_source_refuted.sh +176 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_regression_runner_lanes.sh +101 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_root_depth.sh +6 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +19 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/e2e-real-flow-testing.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +3 -3
- package/dist/assets/release.json +63 -33
- package/package.json +1 -1
|
@@ -59,6 +59,21 @@ Use a background writer fed by a queue with explicit control messages:
|
|
|
59
59
|
or a transcript consumer that a turn/session completed, flush the log. A "done" that outraces the
|
|
60
60
|
writer can lose the last turn on crash. (This mirrors the background-task finality rule in
|
|
61
61
|
`terminal-cli-dev`: retain final output until consumers have acknowledged it.)
|
|
62
|
+
- **Checkpoint before external commitments, not only before finality.** Flush the recorded log
|
|
63
|
+
prefix durable *before a model adapter receives a request*, *before any tool call may produce an
|
|
64
|
+
external side effect — nested and delegated calls (code-mode sub-calls, sub-agent tool use)
|
|
65
|
+
included, since they route through the same effect surface*, and at each pre-step boundary — and treat a rejected/failed checkpoint
|
|
66
|
+
as blocking that dispatch or side effect (fail closed), not as a warning. Finality-only flushing
|
|
67
|
+
leaves a mid-turn window where a crash loses the prior response and ordered tool results that an
|
|
68
|
+
already-executed side effect or the next request acted on. The flushed prefix alone still cannot
|
|
69
|
+
distinguish *never dispatched* from *dispatched, outcome unknown* after a crash — so pair the
|
|
70
|
+
checkpoint with a durable dispatch-intent/attempt record appended-and-flushed before the action
|
|
71
|
+
and a completion record after: for model requests that is exactly §3's attempt-state rule
|
|
72
|
+
(prepared → dispatch-attempted → accepted/uncertain/rejected, idempotency-keyed); apply the same
|
|
73
|
+
shape to non-idempotent tool side effects, where an unresolved attempt is reconciled or surfaced
|
|
74
|
+
as a typed indeterminate outcome — never auto-retried (the task-level counterpart lives in
|
|
75
|
+
`agent-task-orchestration.md`'s effect-finality rule). (This is the event-log half; the §3
|
|
76
|
+
accounting records carry their own commit-before-dispatch rule — complementary, not duplicates.)
|
|
62
77
|
- **Make writer failure observable.** If the writer task dies (disk full, IO error), record a
|
|
63
78
|
terminal-failure state that every later append/flush call surfaces — never let the recorder keep
|
|
64
79
|
accepting events into the void. A dropped persist is a data-loss bug, not a best-effort log line.
|
|
@@ -323,7 +338,42 @@ When a session's estimated context approaches the model window, compact it.
|
|
|
323
338
|
baseline and a window ordinal so "how much have we grown" is measured since the previous compaction,
|
|
324
339
|
not from zero. **Prefer server-reported token usage over local estimation** when the provider
|
|
325
340
|
returns it; estimation is the fallback that drives the trigger when no server count is available
|
|
326
|
-
yet.
|
|
341
|
+
yet. Feed the trigger from **one unified token meter**, not per-consumer recounts: a single
|
|
342
|
+
replay-aware event source (rebuilt from the log on resume/replay, per §4's restore-accounting
|
|
343
|
+
rule), deduplicating repeated *ingestion of the same attempt* (a replayed or re-read response)
|
|
344
|
+
by attempt/request id. Keep two projections over that one source, and derive each correctly:
|
|
345
|
+
the **logical/pressure projection** (compaction trigger, budget displays) measures the *current
|
|
346
|
+
model-visible envelope*: the latest authoritative envelope measurement — server-reported context
|
|
347
|
+
size bound to the *accepted canonical attempt*, not merely the newest report — **plus the
|
|
348
|
+
estimated size of every model-visible item appended since that measurement** (a large tool
|
|
349
|
+
result landing after the report otherwise dispatches an over-window request unpruned), or a
|
|
350
|
+
fresh full-envelope estimate taken before dispatch; never a sum across sequential requests,
|
|
351
|
+
which re-counts the resent history every turn and trips the watermark while the real context
|
|
352
|
+
still fits; the
|
|
353
|
+
**per-attempt/billing projection** sums every real provider attempt — a genuine retry consumes
|
|
354
|
+
and may bill tokens even when its payload duplicates the first attempt. Collapsing the two
|
|
355
|
+
either underreports spend or triggers needless lossy compaction; consumers that recount for
|
|
356
|
+
themselves disagree about when pressure exists.
|
|
357
|
+
- **Run a deterministic pruning layer before the summarization layer.** Before invoking any
|
|
358
|
+
model-written summary, apply a model-free, replay-safe pruning pass over the candidate window —
|
|
359
|
+
dropping or trimming stale tool-call/tool-output payloads under the same keep/drop rules — and
|
|
360
|
+
only summarize if the window is still over pressure afterwards. Pruning is deterministic, cheap,
|
|
361
|
+
and reversible in design terms; jumping straight to LLM summarization pays nondeterminism and
|
|
362
|
+
fidelity loss for reduction that pruning could have achieved. Declare which of two shapes the
|
|
363
|
+
pruning layer is, because a prune that relieves pressure without a summary has no summary-bearing
|
|
364
|
+
compaction marker to ride on: **ephemeral** pruning is applied at context assembly from the same
|
|
365
|
+
versioned rules every time, never mutates the log, and replay reproduces it by re-running the
|
|
366
|
+
rules; **committed** pruning persists a typed prune-only marker — covered range, rule/version,
|
|
367
|
+
resulting kept-set references, and the window-baseline update, with the summary field legitimately
|
|
368
|
+
absent — so resume neither reconstructs the unpruned envelope nor re-triggers compaction. An
|
|
369
|
+
unclassified prune that changes the model-visible envelope without either contract breaks
|
|
370
|
+
faithful replay. (The trim-to-fit rule below stays as the last-resort fallback *after* a
|
|
371
|
+
summary; this layer runs *first*.)
|
|
372
|
+
- **Manual/user-invoked compaction fails with typed codes, not a generic error.** Distinguish at
|
|
373
|
+
least: another compaction already in flight, cancelled, context changed since the request,
|
|
374
|
+
summarization failed, and commit/persistence failed — the caller's correct reaction (retry,
|
|
375
|
+
re-read, give up, surface data-loss risk) is different for each, and a generic failure trains
|
|
376
|
+
users to spam retry across all of them.
|
|
327
377
|
- **Two strategies, chosen by provider capability:** *local/inline* compaction (you send the history
|
|
328
378
|
to the model with a dedicated summarization prompt and replace it with the returned summary) or
|
|
329
379
|
*provider-remote* compaction (the provider compacts server-side). Pick per provider; do not assume
|
|
@@ -361,7 +411,9 @@ When a session's estimated context approaches the model window, compact it.
|
|
|
361
411
|
- Crash-safety requires the §2 preconditions: one exclusive writer per session, whole-record atomic
|
|
362
412
|
append, `fsync` before acking Flush, and a monotonic sequence number.
|
|
363
413
|
- Flush the log before signaling turn/session finality; never let a completion signal outrace the
|
|
364
|
-
writer.
|
|
414
|
+
writer. The same checkpoint discipline gates external commitments mid-turn: flush the recorded
|
|
415
|
+
prefix before model dispatch and before a tool's external side effect, and a failed checkpoint
|
|
416
|
+
blocks the action, fail closed.
|
|
365
417
|
- A failed persist/flush is raised and made observable, never swallowed as empty success.
|
|
366
418
|
- Always persist the structural markers (session meta, turn context, compaction) that replay needs.
|
|
367
419
|
- Truncating a persisted payload is allowed only if it matches what the model saw (or the full
|
|
@@ -51,6 +51,17 @@ When the model requests multiple tool calls in one response and the model/runtim
|
|
|
51
51
|
- Apply post-tool hooks; let them inject additional context, treated as untrusted (see `agent-lifecycle-hooks.md`).
|
|
52
52
|
- Emit per-call telemetry (tool name, decision source, duration, outcome) keyed by call id.
|
|
53
53
|
- Bound result size fed back to the model; truncate large tool output **visibly and structure-aware** — truncating in the middle of structured (JSON) output yields unparseable downstream content. Use structure-aware markers or summarize, don't byte-chop.
|
|
54
|
+
- When bounding **replaces** an oversized plain-text result with a persisted artifact (spill), split the capability per `agent-capability-composition.md` §1: a **spill-store seam** owns persistence only (save verbatim, return an opaque locator plus retrieval hint; reject on real storage failure), swappable providers own the backing store, and a **policy plugin** owns *when* to replace — the model-visible replacement is a bounded head/tail preview plus the locator and an explicit omitted-byte count, never a silent cut. Three semantics make this safe:
|
|
55
|
+
- **Spill is fail-open for the call it wraps, but never for the bound**: a spill failure (no store, save rejected) never converts a successful tool call into an error result — and it also never returns the unbounded original, which would defeat the model-context bound above; the fallback is a clearly-labeled bounded truncation ("full result unavailable", omitted bytes counted). (Persistence-failure *labeling* discipline: §Persisted tool-output artifacts below.)
|
|
56
|
+
- **Exclude read-back tools from model-facing replacement**, or the model reads a file, gets a spilled preview pointing at a file, and reads again — a read→spill→read loop. A durable-log copy of the same output is not model context, so bound it on its own arm with the same cap; the loop exclusion does not apply there.
|
|
57
|
+
- Replacement operates on the **final formatted model-facing result**, not the tool's internal/canonical value — provider-side caps stay mandatory and separate, and the programmatic value consumed downstream is preserved unchanged.
|
|
58
|
+
|
|
59
|
+
## Composable execution guards
|
|
60
|
+
|
|
61
|
+
Cross-cutting per-call guards compose as wrappers around dispatch rather than being woven into the router — but the two below sit in different risk classes, and only one is removable (the absent-plugin default rule of `agent-capability-composition.md` §3 applies):
|
|
62
|
+
|
|
63
|
+
- **Repeat-call reminder** (advisory, removable): detect an identical tool call repeated within the turn/window and inject an advisory notice into the result path so the model can break the loop. Advisory only — it nudges, and must never be relied on as a denial boundary (policy/sandbox layers own denial).
|
|
64
|
+
- **Cooperative timeout** (mandatory enforcement; only the packaging is pluggable): the tool declares its timeout budget and the contract that it honors the abort signal; the wrapper owns the deadline, returns a **typed** timeout failure the model can distinguish from an execution error, and never abandons the still-running work — cancel **and join** it (same discipline as Parallel execution above), or the "timed-out" tool keeps mutating state after the model moved on. Deadline enforcement is a dispatch-lifecycle obligation, not an optional nicety: a deployment without the wrapper plugin still owes a bounded default deadline with cancel-and-join (fail closed) in the dispatch/provider path — an unloaded plugin must never mean "tools run unbounded". With nested deadlines (an outer turn budget and an inner per-call budget), attribute the firing deadline correctly: an inner timeout is that call's typed failure, not a turn abort, and vice versa.
|
|
54
65
|
|
|
55
66
|
## Code-mode: tools as a code/exec runtime (optional advanced pattern)
|
|
56
67
|
|
|
@@ -69,8 +69,15 @@ The fitness function may be a scalar metric, unit tests / validators / schema ch
|
|
|
69
69
|
|
|
70
70
|
Every material model or prompt change should define:
|
|
71
71
|
|
|
72
|
+
- the decision the eval supports, named before choosing datasets or scorers — this reference's working classes: version/model selection (benchmarking), release/regression gate, capability/unit check, and production monitoring and failure diagnosis (informed by the public offline-vs-online eval-type split per `docs.langchain.com/langsmith/evaluation-types` and success-criteria-first design per `docs.anthropic.com/en/docs/test-and-evaluate/define-success`). Conclusions do not transfer automatically across decision classes — each intended decision must satisfy its own predeclared criteria: a selection benchmark's average is not by itself a release verdict (a release verdict needs regression comparison against the incumbent baseline plus red-line blockers — see Eval Reliability and Freeze Comparison Criteria below), and a monitoring signal is not a version verdict; one combined, predeclared design may serve several decisions when every decision's criteria are declared up front and each is satisfied on its own terms (per the combined-design rule under Freeze Comparison Criteria);
|
|
72
73
|
- dataset or replay source;
|
|
73
74
|
- expected output, rubric, or comparator;
|
|
75
|
+
- run protocol: a declared repetition and uncertainty protocol matched to the decision and to the observed nondeterminism —
|
|
76
|
+
- one run per task only when the evaluated path is deterministic end to end, generation included (a deterministic code grader over stochastic model output does not qualify), or with recorded justification that output variance cannot move the decision;
|
|
77
|
+
- where variance can move an offline or replay-based comparison, resample the same task several times and use question-level averages for aggregate quality metrics (question-level-averaging technique per `arxiv.org/abs/2411.00640`, stated there for chain-of-thought evals; applying it to other variance-sensitive comparisons is this reference's extension), while red-line/blocker checks aggregate conservatively (any hit fails, never averaged away);
|
|
78
|
+
- online production-monitoring decisions observe live traffic through predeclared observation windows with minimum volume, persistence, and hysteresis (per the operational-gate rule under Freeze Comparison Criteria and `inference-capacity-operations.md`);
|
|
79
|
+
- failure-diagnosis decisions reproduce the failure under control: sanitized and isolated replay of recorded inputs, with side effects suppressed or stubbed wherever mutation is possible (reproduction discipline per `defect-diagnosis`); live side-effecting requests are never re-executed;
|
|
80
|
+
- for compared runs, hold the conditions fixed — everything except the declared variable under test, with both values of that variable recorded (model string, prompt version, retrieval index, tool set, and decoding parameters are the usual fixed set);
|
|
74
81
|
- quality metrics such as accuracy, recall, consistency, parse success, or groundedness;
|
|
75
82
|
- runtime metrics such as latency, token usage, success rate, fallback rate, and cost;
|
|
76
83
|
- regression examples and human review notes when judgment is subjective.
|
|
@@ -161,3 +161,9 @@ Rollback path post-promotion:
|
|
|
161
161
|
- Force a canary abort → traffic returns to stable within seconds (mesh weight propagation time).
|
|
162
162
|
- Trigger header-lane canary with the opt-in header → request lands on canary subset; without header → lands on stable.
|
|
163
163
|
- Mirror enabled → mirror target sees traffic; primary handler is unaffected; metrics tagged distinctly.
|
|
164
|
+
|
|
165
|
+
## Topic-extension backlog
|
|
166
|
+
|
|
167
|
+
Entries here are registered candidates, not adopted guidance. Each names the candidate, its evidence status, and the condition that unblocks adoption; the round that evaluates one records keep/narrow/discard against its entry.
|
|
168
|
+
|
|
169
|
+
- **Application-layer rolling provider transition (deferred candidate).** Observed form, from one agent-native product repository (evolving portfolio): a new backend provider registers with an in-process broker as an additional route, traffic shifts to it by weight, and the old provider unloads once it has no in-flight work — demoting blue-green from an infrastructure operation to an in-process composition pattern. Evidence status: hypothesis-grade — single source, and the source itself labels the account observational. Do not land this as executable rollout guidance until at least two independent external sources corroborate the pattern in production use; re-evaluate at the next design round touching provider transition, blue-green, or in-process traffic shifting.
|
|
@@ -10,7 +10,8 @@ These are external quality benchmarks, not visual style sources. Do not copy ano
|
|
|
10
10
|
- W3C WCAG 2.2: accessibility criteria for keyboard, focus, contrast, labels, target size, error identification, and consistent navigation.
|
|
11
11
|
- web.dev Core Web Vitals: Largest Contentful Paint, Cumulative Layout Shift, Interaction to Next Paint, and field/lab measurement.
|
|
12
12
|
- Google HEART framework: product-experience metrics across happiness, engagement, adoption, retention, and task success.
|
|
13
|
-
-
|
|
13
|
+
- Apple Human Interface Guidelines (developer.apple.com/design/human-interface-guidelines) and Material Design 3 (m3.material.io): first-party platform specifications for feedback, loading, interaction states, accessibility, and platform conventions — distilled into the Platform Convention Walkthrough section below. Each criterion there is labeled (HIG) or (Material); a single-source criterion is that platform's convention, not a cross-platform standard.
|
|
14
|
+
- Other design-system guidance such as Atlassian Design System: secondary confirmation for touch targets, empty/error states, progressive disclosure, feedback, and content clarity.
|
|
14
15
|
|
|
15
16
|
## Heuristic Review Layer
|
|
16
17
|
|
|
@@ -38,6 +39,54 @@ Before launch or review, verify:
|
|
|
38
39
|
- **Reduced-motion / color-scheme / contrast / transparency design intent must be declared, not auto-derived from `@media` query alone**. For each non-decorative animation, name whether it is *essential* (progress indicator, drag preview, view transition that conveys a state change) or *decorative* (parallax, autoplay carousel, hover bounce); `prefers-reduced-motion: reduce` should remove or replace decorative motion by default and may shorten essential motion but cannot omit feedback. An explicit in-product opt-in (e.g. user setting "Show celebration animation even when system asks for reduced motion") MAY override the default for brand-splash / completion-celebration moments when policy permits, but the override must be opt-in not opt-out and the default behavior must respect the system preference. `prefers-color-scheme` requires the design system to ship Light + Dark token pairs (no missing pair = no dark-mode claim). `prefers-contrast: more` and `prefers-reduced-transparency` are useful supplements where browser support permits but are NOT Baseline yet — do not gate accessibility compliance on them; route them through a separate "increased contrast" theme variant when product needs require it.
|
|
39
40
|
- **APCA (Accessible Perceptual Contrast Algorithm) is the WCAG 3 / Silver candidate contrast method, not a WCAG 2.2 replacement**. WCAG 2.2 SC 1.4.3 / 1.4.11 (4.5:1 body / 3:1 large text + UI components) remains the legal/audit baseline. APCA can be used as a *supplementary* perceptual check (per APCA Bronze Simple Mode, Lc 75 minimum / Lc 90 preferred for body text) when WCAG 2.x mathematical contrast passes but the result looks washed-out, or when designing dark mode where WCAG 2.x ratios systematically over-permit low-readability combinations. Decision matrix: WCAG 2.x fail blocks accessibility/legal compliance claims regardless of APCA result; WCAG 2.x pass + APCA fail is NOT a WCAG failure but should be treated as a readability / product-quality defect — either adjust the color tokens to also pass APCA Bronze, or document the acceptance with a rationale (brand constraint, dark-mode literal preserved). Do not ship a design that passes only APCA but fails WCAG 2.2.
|
|
40
41
|
|
|
42
|
+
## Platform Convention Walkthrough (HIG / Material)
|
|
43
|
+
|
|
44
|
+
Use this section when reviewing or accepting a mobile-platform surface (iOS/iPadOS or Android/Material-based, including Flutter/React Native apps that adopt a platform design language). Each criterion is a pass/fail walkthrough check distilled from the first-party spec named in its label; verify against the rendered surface, not the design file alone. These complement — never replace — the WCAG 2.2 acceptance line above: where a platform minimum is stricter than WCAG (e.g. target size), the platform minimum is the walkthrough bar for that platform. Criteria mirror each source's own normative strength: where the source states a recommendation ("ideally", "consider", "in general", "in most cases"), a recorded, justified exception passes as an exception — only a silent shortfall fails; where the source states a requirement, the failure is unconditional.
|
|
45
|
+
|
|
46
|
+
### State completeness against the platform spec
|
|
47
|
+
|
|
48
|
+
- **Interaction-state matrix is complete for the component class and input modalities (Material).** Every interactive component accounts for each state its Material component class inherits — enabled, plus disabled/hover/focused/pressed/dragged only where the class takes them (action/selection/input components inherit most; app bars, dialogs, menus, navigation components inherit few) — across the input modalities the surface ships on (hover needs a pointer; focused needs a focus-capable input such as keyboard or voice). States the class does not inherit are marked inapplicable, never styled in. Fail: an action component on a keyboard-capable surface styles only enabled and pressed.
|
|
49
|
+
- **Disabled semantics are real, not painted (Material).** A disabled component cannot be focused, dragged, or pressed, and does not change state when tapped or hovered — unrelated explanatory feedback (a tooltip saying why it is disabled) is not prohibited. Components whose class does not take a disabled state in Material — app bars, badges, dialogs, FABs, menus, navigation bar/drawer/rail, sheets, tabs, tooltips — never render a "disabled" look: when a FAB's action is unavailable, remove the FAB rather than disabling it. Fail: a grayed-out FAB or tab sits on screen, or a "disabled" card still accepts a drag.
|
|
50
|
+
- **Each state change is signaled by more than one visual cue (Material).** Material's baseline is two visual indicators per state so state remains perceivable under color-vision or contrast loss; opacity-only or color-only state styling fails.
|
|
51
|
+
- **Transient input states are singletons (Material).** At most one hover, one focus, one pressed, and one dragged state visible at a time in a layout; persistent states (selected, activated) may combine with them on the same element (e.g. a selected chip showing hover).
|
|
52
|
+
- **Feedback reaches people through more than one channel (HIG).** Significant feedback pairs color with text/icon, and sound with haptic where sound is used, so it survives a silenced device, a glance away, or a screen reader. Fail: success/failure conveyed by hue change alone or by sound alone.
|
|
53
|
+
- **Interruption level matches significance (HIG).** Passive status renders in-context near the item it describes (badge, inline line); modal alerts are reserved for critical, ideally actionable information. Fail: routine status delivered as a modal, or a data-loss warning delivered as a passive toast.
|
|
54
|
+
- **Data-loss warnings fire on the unexpected-and-irreversible boundary, both directions (HIG).** Warn before an action whose data loss is unexpected and irreversible; do NOT interpose confirmation when loss is the expected result of the user's own action (e.g. moving a file to trash). Fail in either direction: silent irreversible loss, or confirmation nagging on expected outcomes.
|
|
55
|
+
- **Completion feedback is reserved for significant outcomes; failure feedback is never omitted (HIG).** People expect success, so confirm only payment-grade/significant completions — but every command that cannot be carried out must say so and say why, with the next step. Fail: a no-op button press with no explanation.
|
|
56
|
+
- **Content loading shows something immediately and frees the user (HIG).** HIG scopes this to content/asset loading: placeholder/skeleton content appears at once instead of a blank wait, loading continues in the background so unrelated safe actions stay available, a determinate indicator is used when duration is known and indeterminate only when it is not, and an unavoidably long load gets meaningful interim content. This does not apply to in-flight mutations (payment, deletion, submission): while one is pending, its duplicate or conflicting mutation controls are blocked per the high-risk resilience states in `SKILL.md`, not left available.
|
|
57
|
+
|
|
58
|
+
### Accessibility against the platform spec
|
|
59
|
+
|
|
60
|
+
- **Text scales to 200% without breaking the layout (HIG, stated as "ideally").** Support the platform text-size setting (Dynamic Type on Apple platforms) up to 200% enlargement — HIG's recommended target, so a smaller ceiling passes only as a recorded exception; the walkthrough re-renders key screens at enlarged sizes and checks truncation, overlap, and control reachability. Adoption mechanics (Dynamic Type APIs, per-platform text-scaling behavior) belong to the stack implementation owner, not this walkthrough.
|
|
61
|
+
- **iOS hit regions measure at least 44×44pt (HIG, stated as "a general rule").** Every control's hit region is at least 44×44pt (60×60pt on visionOS), and the padded hit area is what must measure up, not the visual glyph; a smaller region passes only as a recorded exception. The ~12pt padding around bezeled elements and ~24pt around bezel-less ones is HIG's "generally works well" guidance, checked the same way. Stricter than WCAG 2.2's 24px floor — the platform benchmark is the walkthrough bar.
|
|
62
|
+
- **Android touch targets measure at least 48×48dp with 8dp spacing (Material, stated as "consider" / "in most cases").** Touch targets at least 48×48dp with at least 8dp between targets, pointer targets at least 44×44dp, the padded target extending beyond the visual bounds; a shortfall passes only as a recorded exception. Stricter than WCAG 2.2's 24px floor — the platform benchmark is the walkthrough bar.
|
|
63
|
+
- **Contrast meets the W3C-derived platform bar (Material, citing W3C).** Small text at least 4.5:1 against background; large text (14pt bold / 18pt regular and up) and meaningful graphics at least 3:1. Clustered non-text containers (e.g. a button group) need 3:1 container-vs-background. The standalone-prominence exemption (a FAB) applies only to that container-vs-background ratio — the element's own text, icons, focus indicators, and meaningful graphics still need their 3:1/4.5:1 bars; disabled states are the only class exempt from contrast requirements.
|
|
64
|
+
- **Reading and focus order follows content hierarchy (Material).** Screen-reader order follows the top-down source/DOM structure, headings do not skip levels, one H1 per web page, repeated landmarks get unique labels. Fail: visual order diverges from traversal order with no remediation.
|
|
65
|
+
- **Focus is managed across context changes (Material).** Initial focus is defined per screen; opening a dialog moves focus into it; closing returns focus to the element that opened it; a visible focus ring appears on keyboard traversal. Fail: focus lost to page top after a dialog closes.
|
|
66
|
+
- **Labels describe purpose, not appearance, and omit the role (Material).** Icon-only controls, meaningful images, and progress/error cues carry labels naming the action or meaning ("Voice search", not "Microphone"); decorative images are hidden from assistive tech; the role word ("button") never appears inside the label.
|
|
67
|
+
- **Core functionality is never gesture-only (HIG).** Any action in the UI's core functionality or supported task flows that a gesture performs (swipe-to-dismiss, swipe-row actions, custom gestures) is also reachable through a visible onscreen control; an optional convenience gesture duplicating an already-visible control needs no second alternative. Frequent actions use the simplest gesture available, no custom multi-finger requirements.
|
|
68
|
+
- **Keyboard access is complete and system shortcuts stay untouched (HIG).** Core flows complete with the keyboard alone (Full Keyboard Access on Apple platforms), and system-defined keyboard shortcuts are not overridden.
|
|
69
|
+
- **Custom shortcuts default to two-key combinations and are discoverable (Material).** Custom keyboard shortcuts use two or more keys by default — a single-key shortcut needs a remap option, component-focus scoping, or an off switch — and a help surface lists them.
|
|
70
|
+
- **Timed UI does not self-dismiss content people must act on (HIG).** Views and controls that auto-dismiss on a timer are minimized; anything carrying a decision or unfinished reading dismisses by explicit action. Fail: an error toast that disappears before its recovery action can be reached.
|
|
71
|
+
- **Reduce Motion is honored with concrete substitutions (HIG).** When the OS reduce-motion setting is on, decorative/repetitive animation stops by default — the Accessibility Baseline's narrowly scoped, explicitly opt-in in-product override for brand-splash/celebration moments remains valid and is the only exception. For animations that use these effects: springs tighten (no bounce), x/y/z transitions become fades, z-depth and blur animations are avoided, and gesture-driven animation tracks the gesture. Essential status motion (progress) remains. Declare essential-vs-decorative intent per the reduced-motion rule in the Accessibility Baseline above; native OS setting detection routes to `platform-mobile-patterns.md` Mobile Motion Discipline.
|
|
72
|
+
|
|
73
|
+
### Platform conventions walkthrough
|
|
74
|
+
|
|
75
|
+
- **One-handed reachability shapes iPhone layouts (HIG).** Primary and frequent controls live in the middle or bottom of the screen; back-swipe from the edge and list-row swipe actions are preserved, not hijacked by custom edge gestures (Android's equivalent predictive-back geometry routes to `platform-mobile-patterns.md`).
|
|
76
|
+
- **The surface adapts to user-chosen appearance settings (HIG).** A screen passes walkthrough only after being checked under orientation change, Dark Mode, and enlarged Dynamic Type — the platform treats these as user choices the app must follow, not edge cases.
|
|
77
|
+
- **Standard platform components are the default for standard tasks (Material).** Standard platform controls and semantic elements inherit assistive-technology support for free; a custom replacement for a standard task (e.g. a non-standard dialog) carries the burden of extra AT verification before it passes walkthrough.
|
|
78
|
+
- **Contrast and appearance adaptation is verified on the rendered surface (HIG).** Beyond the orientation/Dark Mode/Dynamic Type checks above, the walkthrough gate is observable: the surface renders correctly with the Increase Contrast setting on — text, icons, and state indicators keep sufficient contrast and their meaning. Preferring system-defined colors (whose accessible variants adapt automatically) and familiar system behaviors is non-blocking implementation guidance verified in code review, because two token implementations can render identically.
|
|
79
|
+
|
|
80
|
+
### Deliberately not absorbed from the platform specs
|
|
81
|
+
|
|
82
|
+
Recorded so future rounds do not re-import them; each was read and rejected for walkthrough use:
|
|
83
|
+
|
|
84
|
+
- HIG media-accessibility taxonomy (captions vs subtitles vs audio descriptions vs transcripts) — media-content-type guidance, not a screen-walkthrough criterion; consult the HIG Hearing section directly when shipping media surfaces.
|
|
85
|
+
- HIG platform-capability integrations (Siri/Shortcuts, Switch Control, Voice Control setup, Assistive Access optimization) and watchOS/visionOS-specific rules — implementation- or platform-mode-specific; route to the stack implementation skill if those surfaces enter scope. One exception is retained: the visionOS 60×60pt hit-region figure stays inside the iOS hit-region criterion as informational context only — this walkthrough's scope remains iOS/Android mobile surfaces and does not govern visionOS.
|
|
86
|
+
- Material state-layer token mechanics (fixed opacity percentages, on-color derivation) — design-kit implementation detail owned by token/component references, not an acceptance criterion.
|
|
87
|
+
- Material web-landmark role enumeration (the eight ARIA roles) — imported only as the "landmarks get unique labels" criterion; the full role catalog is reference material, not a checklist.
|
|
88
|
+
- Visual-style content from either spec (Liquid Glass materials, M3 Expressive shapes/motion values) — style adoption is a product decision covered by `platform-mobile-patterns.md` Platform OS Updates; copying platform visual language is already ruled out by "What Not To Absorb" below.
|
|
89
|
+
|
|
41
90
|
## Performance And Perceived Speed
|
|
42
91
|
|
|
43
92
|
Use these checks for AI, feed, media, upload, and document-heavy surfaces:
|
|
@@ -45,6 +45,7 @@ Always include concrete file/line references for code reviews and Figma file/pag
|
|
|
45
45
|
7. **Check accessibility basics**: readable contrast, keyboard/focus where relevant, touch target size, visible labels, reduced-motion risk, safe-area/keyboard behavior on mobile.
|
|
46
46
|
8. **Check visual craft**: anti-slop, product-level identity, spacing rhythm, typography scale, consistent iconography, appropriate density.
|
|
47
47
|
9. **Check serious-domain adaptation** when relevant: source, timestamp, partial data, confirmation, audit labels, and no unsafe optimistic UI.
|
|
48
|
+
10. **Check platform-convention conformance** for iOS/Android app surfaces: run the pass/fail criteria in `external-ui-ux-quality-benchmarks.md` Platform Convention Walkthrough (HIG/Material state completeness, platform accessibility minima, and platform conventions) against the rendered surface.
|
|
48
49
|
|
|
49
50
|
## UI Checks
|
|
50
51
|
|
|
@@ -12,6 +12,7 @@ Use this as the first reference for Python backend, Python microservice, AI-serv
|
|
|
12
12
|
|
|
13
13
|
- Decide the backend shape first: API service, internal service, worker service, scheduled job service, SDK/package support library, CLI, AI-service host, or modular monolith.
|
|
14
14
|
- Use separate deployable service boundaries only when ownership, scaling, data ownership, runtime isolation, deployment cadence, or rollback needs justify them. Keep one modular service/package/script shape when boundaries are unclear or the scope is too small for separate services.
|
|
15
|
+
- External grounding for the boundary rule above (adopted in part): [Bounded Context](https://martinfowler.com/bliki/BoundedContext.html) (Fowler's overview of the central DDD strategic-design pattern; origin credited there to Eric Evans, *Domain-Driven Design*) — a boundary is where one internally consistent model stops being valid, so boundaries follow model and responsibility lines, never noun count — and [Conway's Law](https://martinfowler.com/bliki/ConwaysLaw.html) (Melvin Conway, ["How Do Committees Invent?", 1968](https://www.melconway.com/Home/Committees_Paper.html); Fowler's overview) — a system's structure mirrors the builders' communication structure, which is why team ownership and communication structure are a legitimate split input. Two limits: the ownership, scaling, release-cadence, and rollback tests in the rule above are this skill's own operational criteria, attributed to neither source; and a bounded context does not require its own deployable service — the modular-monolith default stands, so these sources justify model/ownership boundaries, not a service-per-context split. Borrowed scope: boundary-input criteria only; no claim to the full DDD strategic-design method (context maps, ubiquitous language) or an inverse-Conway process. Sibling: `go-microservice-architecture/references/architecture-playbook.md` ("Service Boundary Rules") carries the same grounding; keep the two in sync.
|
|
15
16
|
- For each Python microservice that is justified, define the owner, contract, data source of truth, inter-service auth, timeout/retry/fallback policy, deployment unit, canary/rollback, and observability.
|
|
16
17
|
- Separate layers:
|
|
17
18
|
- transport/framework: routing, auth, validation, response mapping
|
|
@@ -18,3 +18,9 @@ Use this for pyproject, uv/poetry/pip, lockfiles, tooling, containers, and deplo
|
|
|
18
18
|
- Release readiness includes migrations, startup validation, smoke tests, canary, rollback, and observability checks.
|
|
19
19
|
- **ASGI server choice has expanded beyond `uvicorn` / `gunicorn+uvicorn` / `hypercorn`** — `granian` (emmett-framework, Rust-based) is the current credible high-throughput alternative for ASGI services, supporting ASGI/3, RSGI, WSGI, HTTP/1, HTTP/2, TLS, WebSockets (HTTP/3 planned per the project README's "eventually 3" roadmap note — verified not shipped as of May 2026). Per the Granian project's own `benchmarks/vs.md` and third-party load-test repos (e.g., `piccolo-orm/asgi_server_performance`, `synodriver/asgi-server-benchmark`), ASGI echo on 10KB payload typically lands granian > uvicorn-httptools > hypercorn by roughly the ratios 58k / 51k / 8k RPS in April-2026-era runs; file-serving gap is wider (granian ~47k vs uvicorn ~18k via `pathsend`). Treat the absolute numbers as benchmark-snapshot-specific; re-run against your workload before basing a switch on them. Architecture impact: when serving throughput is the binding constraint, granian can buy headroom without rewriting the **plain ASGI path**. **"Without rewriting" caveats**: granian's worker / process model differs from `gunicorn+uvicorn` fork-based workers (Rust runtime + Python interpreters with different lifecycle hooks); ASGI lifespan events, contextvars propagation across worker boundaries, custom signal handlers, prometheus/metrics exporters tied to uvicorn internals, and ASGI middleware that depends on uvicorn-specific behavior all need smoke-testing on granian before a switch. If the team plans to adopt granian's bespoke RSGI protocol for max performance (rather than ASGI), application code that uses ASGI-specific middleware, ASGI scope manipulation, or third-party ASGI libraries WILL need rewriting — RSGI is a different protocol, not a faster ASGI. Trade-offs: smaller operational maturity, fewer community recipes, Rust-runtime-on-the-side observability differs from a pure-Python server. **Choose uvicorn** for ecosystem maturity, broad reference material, and known operational patterns; **choose granian** when (a) profiled benchmarks on your workload show uvicorn saturation, (b) the team has Rust-toolchain debugging capacity, (c) the deployment story can absorb a less-common runtime. Hypercorn remains the choice when HTTP/2 + ASGI under pure-Python ops matters more than peak throughput.
|
|
20
20
|
- **Python runtime version baseline (2025-2026)**: Python 3.13 (released October 2024) ships **experimental** free-threaded build per PEP 703 — GIL-disabled, ~40% single-threaded perf hit per python.org "What's New in 3.13" notes, used at the team's risk for parallel-CPU workloads. Python 3.14 (released 7 October 2025) advances free-threading to **supported (Phase II of PEP 703)** per PEP 779 — meaning the free-threaded build is a first-class supported configuration, NOT that it is the default Python build or the default production choice. Per the python.org free-threading howto, the single-threaded penalty narrowed to ~5-10% (specializing adaptive interpreter re-enabled thread-safely); PEP 803 defines the `abi3t` stable ABI for free-threaded C extensions. Architecture impact: for services where parallel CPU work matters (in-process ML inference fan-out, heavy parsing, parallel compression), 3.14 free-threading is the first version where adoption is **a defensible experiment for a selected service**, not yet a defensible default. Pre-flight burn-in required before any production switch: (a) C-extension readiness — all extensions in the service's dependency tree must declare free-threading support; many popular extensions (numpy, pandas, lxml, psycopg native bits, asyncpg native bits, pillow, cryptography) were still mid-migration at 2026-Q1, verify per-version; mixing GIL-only and free-threading-aware extensions in one process is unsupported; (b) GC and runtime behavior under sustained threading load differs from GIL build — measure tail latency, memory residency, and CPU efficiency on the actual workload; (c) debugger / profiler ergonomics (gdb, pdb, py-spy, scalene, prometheus exporters) may have rough edges on free-threaded builds; (d) library-level thread-safety: code paths that were "implicitly safe because of the GIL" can race in free-threaded mode (singletons built at import time, module-level mutable caches, third-party libraries that rely on GIL-protected dict mutation). For pure I/O-bound services, stay on stock GIL build — the 5-10% overhead is pure cost. For services on 3.13 or older, treat free-threading as opt-in research, not default.
|
|
21
|
+
|
|
22
|
+
## Topic-extension backlog
|
|
23
|
+
|
|
24
|
+
Entries here are registered candidates, not adopted guidance. Each names the candidate, its evidence status, and the condition that unblocks adoption; the round that evaluates one records keep/narrow/discard against its entry.
|
|
25
|
+
|
|
26
|
+
- **Package-owned runtime-invariant registries (weak-keep candidate).** Observed form, from one agent-native product repository (evolving portfolio): each package contributes its own invariant checks from a companion module, normal entrypoints do not import the diagnostics layer, and allowlist/blocklist configuration selects which checks run. Evidence status: weak — single source; generalization and owner placement are undecided. Evaluate at the next Python architecture round touching runtime readiness, startup validation, or diagnostics, and record the outcome against this entry. The Go-stack sibling registration lives in `../go-microservice-architecture/references/cross-cutting-concerns.md`.
|
package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md
CHANGED
|
@@ -160,7 +160,7 @@ Use this skill to turn observed experience into durable agent skills without cop
|
|
|
160
160
|
- The trigger is a correction about *reusable skill/process behavior*, NOT every bug/QA/review nit handled inside its own owner skill. Do not wait for the user to say "沉淀": if the user points out a missed source, missed sibling skill, shallow rule, overclaim, domain leakage, missing trigger, missing verification, repeated correction, **or that this workflow should have been invoked at all (an under-trigger / "should you have used 提炼/复盘" correction, including outside an active extraction)**, run correction RCA, update the smallest owning skill/reference/validator, and verify the prevention point before finalizing the turn.
|
|
161
161
|
- When the user asks whether a lesson was durably landed after a failed extraction, verify the actual skill diff or file content first. Do not answer from memory or intent. If the prevention rule is not present in the owning skill, add it or state that it has not been durably landed.
|
|
162
162
|
- **Consolidate and retire rules; a skill's rule set must not grow monotonically.** Every correction adds a guard, but an N-bullet wall on one theme is itself the over-prescription/unreadability failure, and "just append another bullet" is how it regrows.
|
|
163
|
-
- **
|
|
163
|
+
- **Prose rules compete for a finite attention budget, and joint satisfaction degrades with the number of constraints.** (Predecessor claim — only the most-salient rule applies, the rest dormant — **withdrawn**, unsupported.) **Descriptive, not permissive**: attention limits never excuse a violated rule, and are not a reason to refuse a needed one. (a) **Merge-into-canonical beats append**: appending adds contradiction surface and spends budget. (b) Do not rely on co-resident prose for requirements that must hold JOINTLY — structure them as a **walked enumeration at their firing point**: walking a finite list works where holding a conjunction does not. It is why "the rule was loaded" never predicts "the rule was applied". Detail: `references/external-practice-controls.md#instruction-following-mechanisms`.
|
|
164
164
|
- A rule's TEXT in a `SKILL.md`/reference is **living** — merged, tightened, or retired in place — unlike the `source-register.md` ledger (append-only + supersede-by-pointer; never edit/delete a row). "Land a durable prevention point" (per the always-land rule) is satisfied by **a merge into an existing canonical rule, a validator/checklist gate, or a reference pointer — NOT necessarily a new top-level bullet**: before appending a rule, grep the section for an existing owner of the same failure-class and merge instead (this extends the `keep/merge/discard` conflict rule from *incompatible* to *redundant* rules). A consolidation that rewrites or retires rule text is a non-wording shared-skill change whose **behavioral-evidence row (`semantic-control`) MUST carry a zero-loss obligation map** — every before-obligation maps to surviving or genuinely-subsumed text, none silently dropped; the challenge inspects that map. Dropping a real guard under the banner of "consolidation" is a regression.
|
|
165
165
|
- Wording-level dedup mechanics route to `tighten-doc`; the at-add-time consolidation check, the form-by-failure drafting table, the obligation-table format (incl. the verifiable-survivor-pointer rule), the package/support-file integrity axis, strict "subsumed" criteria, register-row boundary, and audit recovery: `references/rule-consolidation.md`.
|
|
166
166
|
|
|
@@ -46,7 +46,7 @@ Disposition:
|
|
|
46
46
|
- Supported: a rule's existence is not proof it executed; use behavioral assertions and coverage/firing evidence. The digest-binding practices these sources describe (complete subject sets, provenance, raw-result digests) are sound for supply-chain trust boundaries where authors and verifiers are distinct parties; the local policy below explains why this repository adopts the firing/coverage principle but not the digest binding.
|
|
47
47
|
- Local evidence policy: the impact-chain gate machine-verifies what is cheap and deterministic — an owner-scoped firing path that resolves to its round's added lines (a unique anchor on a changed normative numbered/list rule, or a changed owner executable), the letters/digits-free wording-only classification computed from that round's owner diff, and the owner-level floor that a non-wording package carries at least one `RED-baseline` row (a `semantic-control` label may supplement but never close a package alone, because an author-selected stable label cannot vouch for a different hidden delta). The `behavioral-evidence` and `observed-failure` fields themselves are required author declarations. A digest-bound attestation apparatus for these rows (in-toto-style subject digests, same-prompt model result pairs, command-result envelopes) was built, evaluated against real iteration, and deliberately removed: under the unsigned-repository-local trust model the author can regenerate every hash, so the apparatus only detected stale records — while costing a full-suite rerun and whole-evidence regeneration whenever any owner script changed by a single byte. That cost defeated normal multi-commit iteration (it broke its own author's branch twice), so behavior claims rest on the firing-path gate, honest authorship, and the mandatory independent review/challenge instead of hashes.
|
|
48
48
|
- Local trust model: register rows remain honest-but-fallible workflow evidence, not a hostile-author security boundary. The gate proves that a changed, normative, owner-scoped rule line (or changed owner executable) exists for every claimed firing path; it does not prove a model run occurred, that a named executable implements the claimed enforcement (a shebang stub passes the static check), that a mangled or ambiguous ledger row was honest (those are warned, not blocked, to avoid false positives on other table shapes), author identity, or non-tampering by an authorized contributor. Independent review/challenge and the fixed checker remain the assurance case.
|
|
49
|
-
- Machine format (relocated from the `SKILL.md` firing-mechanism rule; the local evidence policy above carries the rationale): every added source-register row must carry `behavioral-evidence: RED-baseline` (any observed delta — `observed-failure: yes` requires it) or `semantic-control` (only with `observed-failure: no`), an `observed-failure: yes/no` state, and an owner-scoped `firing-path` — each declaration in its own semicolon-delimited fragment of the cell (`…prose; behavioral-evidence: …; observed-failure: …; firing-path: …`), so a key embedded mid-prose never parses as a declaration. The firing-path anchor is at least 16 characters, occurs once in the file and once in its round's added lines, and lands on a numbered/list Markdown rule with a normative action.
|
|
49
|
+
- Machine format (relocated from the `SKILL.md` firing-mechanism rule; the local evidence policy above carries the rationale): every added source-register row must carry `behavioral-evidence: RED-baseline` (any observed delta — `observed-failure: yes` requires it) or `semantic-control` (only with `observed-failure: no`), an `observed-failure: yes/no` state, and an owner-scoped `firing-path` — each declaration in its own semicolon-delimited fragment of the cell (`…prose; behavioral-evidence: …; observed-failure: …; firing-path: …`), so a key embedded mid-prose never parses as a declaration. The firing-path anchor is at least 16 characters, occurs once in the file and once in its round's added lines, and lands on a numbered/list Markdown rule with a normative action. A row that survives at HEAD must also resolve to an owner this range actually changes — an owner reverted to its base bytes by a rebase or a base-side conflict resolution leaves the changed set while its row stays behind, and the row then vouches for a change the delivered diff does not contain. There is no author-declared escape from this: a corrective rewrite that back-fills a row for a round which merged red produces the same shape, and it is a deliberate, person-adjudicated repair that can adjudicate this refusal too.
|
|
50
50
|
|
|
51
51
|
- **Round scoping — a row is judged against the round it landed in, never the accumulating range.** A row is authored against one round's diff, so reading the whole `base..HEAD` range to classify it judges the row against work it never described. That mismatch produced both directions of the same defect: an already-gated row turned red once a LATER round touched the same owner (which is what the ledger's superseded-row notes were absorbing), and a description-only round lost its routing-surface locator because an EARLIER round had edited that owner's body. The gate cuts rounds at the commits that touch the ledger, walked first-parent so one merged worktree round is one boundary, and each round spans from the previous boundary so work commits sit in the round whose ledger append describes them. The partition is derived from git alone — an author cannot nominate, widen, or move their own scope.
|
|
52
52
|
- Both obligations move together, in opposite directions. **Classification** narrows to the round: whether a diff is wording-only, an identifier retarget, or description-only is asked of that round's bytes, which is what makes a verdict stable once it lands. **Presence** narrows to the round too: the round that changed an owner is the round that owes the row, so owner work committed after a ledger append can no longer ride on an earlier round's row. Narrowing classification without narrowing presence would have opened exactly that laundering route.
|
|
@@ -55,3 +55,61 @@ Disposition:
|
|
|
55
55
|
- Machine-verified no-behavior classes (each drops the firing-path requirement, each requires `observed-failure: no`, and neither is an author waiver — the gate recomputes the predicate from the owner's diff for that row's round and refuses on any mismatch): `not-required wording-only` when no changed line differs by a letter or digit and every changed owner file is Markdown prose; `not-required identifier-rename` when rewriting the base bytes with git-derived rename pairs reproduces the head bytes exactly. Both exist because the anchor uses the SHAPE of a changed line as a proxy for "an obligation changed", and both are diff shapes that carry no obligation at all.
|
|
56
56
|
- Canonical field locator for a routing-surface-only owner — NOT a third no-behavior class. When an owner's entire change is the `SKILL.md` frontmatter `description` entry (same body, same other top-level keys, the value still a YAML string, all checked byte-exact), its firing path is the constant `file:skills/<owner>/SKILL.md#description`. The row still declares `RED-baseline`, still owes its behavioural evidence — for a routing change that is the measured routing delta — and only the free-text anchor is replaced. It is deliberately not an exemption: a description edit decides which requests reach the skill, so exempting it would drop evidence from the class carrying the most behaviour. The locator is canonical rather than a substring because three successive substring designs were refuted — a numbered/list rule shape, then the `description:` line, then any line of the entry — the last because a substring that merely SURVIVES the edit identifies nothing. Same class three times is the signal to drop the proxy, not to patch it again; the predicate is what carries the proof, and the locator only names the field it proved. The predicate matters more here than in the two classes above because a wrong judgement is LOOSER rather than stricter, so it accounts for every changed byte and refuses whatever it cannot account for. A further shape that cannot be located is a signal to replace the proxy, not to widen again. The precondition judges BEHAVIOUR, not file count: the rest of the package — and the entrypoint's own body and other frontmatter keys — may differ by a rename retarget, because that is already a machine-proven no-behaviour class, and the two compose under one normalizer (apply git's rename pairs to the base bytes, then require the head bytes exactly, mode included). Without the composition an owner on an integration branch that accumulates rounds gets refused for retargets it already declared no-behaviour, while having no other rule to anchor on. The description entry is the one part exempt from the normalizer, because it is the change being evidenced; anything the normalizer cannot account for — one real byte — takes the locator away. Siblings are further restricted to regular-file Markdown: a script's bytes reproducing under a substitution says nothing about whether rewriting that identifier was safe there, and a `.md` SYMLINK's blob is its target, so a retargeted target reproduces while what the path resolves to changes. **Disclosed residual, unchanged from the rename class it composes with**: a Markdown sibling can still carry a fenced command or a persisted key whose old identifier is deliberately preserved — an installed artifact name that must not follow a source rename is the observed instance — and rewriting one of those IS a behaviour change the normalizer clears. The exposure is identical to the identifier-rename class, which already reproduces whole packages under the same pairs; this composition inherits that residual rather than widening it, and the pairs still come only from git's own tree state, so a slug is rewritten only when that skill was actually renamed in the same diff. Closing it means distinguishing prose mentions from executable/persisted ones inside Markdown, which needs its own evidence round. Two shapes are knowingly left refused, and refusal here costs nothing that was not already lost: before this locator existed no routing-surface owner could satisfy an anchor at all, so a refused shape simply keeps that prior behaviour rather than regressing from it. (1) A sibling frontmatter value whose YAML tag or class `safe_load` will not construct makes the parse fail, and a failed parse refuses rather than guesses. (2) The owner that ships `references/source-register.md` cannot use the class at all, because its own evidence row lands inside its package and so its diff is never the description alone — a consequence of not special-casing that file, which is what would otherwise let arbitrary ledger content ride along. Both are usability limits on a strictly additive path, not bypasses; closing either one means loosening a predicate whose wrong answers are LOOSER than the bar, so neither is worth doing without its own evidence.
|
|
57
57
|
- Boundary: OPA coverage supports the general distinction between policy presence and policy execution; the repository does not run OPA. in-toto/SLSA supply-chain formats were consulted for the evaluated-and-removed digest apparatus; the shipped gate deliberately does not bind digests and does not reuse or claim their security levels.
|
|
58
|
+
|
|
59
|
+
### `source-refuted` — withdrawing a claim a primary source refutes
|
|
60
|
+
|
|
61
|
+
A rule can fire perfectly and still be false. Removing such a claim usually produces **no measurable behavioural delta** — there is nothing for a `RED-baseline` to capture — while `semantic-control` is refused by the owner-level RED floor. Before this class existed the cheapest path was therefore to *leave the false claim in place*, which is the wrong incentive.
|
|
62
|
+
|
|
63
|
+
The class is not a waiver; it swaps one kind of evidence for another. All four bars are machine-checked and a row fails without any of them:
|
|
64
|
+
|
|
65
|
+
1. `observed-failure: no` — this gate asks whether the owning rule failed to *fire*; a withdrawn claim fired fine, it was simply wrong. Self-declaring `yes` cannot use this class.
|
|
66
|
+
2. The evidence cell cites a **primary source that refutes the claim** (a URL).
|
|
67
|
+
3. The evidence cell carries a pointer to a **zero-loss obligation map** that is a git-tracked regular `.md` inside the repository (no absolute path, no `..`, mode `100644`), whose **anchor resolves to an actual heading** — not merely to a substring, or `#a` would pass on any file containing the letter — and the map **quotes verbatim every substantive line the round deleted**. That last bar is what binds the pointer to the withdrawal itself: to remove a real obligation you must first copy it into the map, where a reviewer sees it.
|
|
68
|
+
4. The owner's round diff is a **pure deletion**, classified structurally rather than by line text: every changed path is a regular `.md` under that owner, every path's status is a modification, **both tree modes are `100644` and identical**, and the diff adds **zero** lines while deleting at least one. Read the modes from `--raw`, not `--name-status`: a `chmod` shows up as a plain modification there and contributes `0/0` to `--numstat`, so a mode change would otherwise ride along with a genuine deletion. Two earlier drafts used proxies — a net-byte floor, then "adds no normative rule line" — and adversarial review broke both: the first by offsetting a smuggled rule with unrelated deletions, the second because a script edit or an `Always`-phrased rule matches no prose predicate. A proxy for the invariant gets bypassed; the invariant itself does not.
|
|
69
|
+
5. **Every** row for that owner in the round is in this class. A lone compliant withdrawal must not lift the floor for a sibling row carrying a real change.
|
|
70
|
+
|
|
71
|
+
Because a withdrawal is a pure deletion it has no added line to anchor on, so this class drops the firing-path requirement exactly as the two `not-required` classes do. Write the explanation of *why* the claim was withdrawn into a reference, not beside the line being removed — an inline note is an added normative line and bar 4 refuses it.
|
|
72
|
+
|
|
73
|
+
These bars are **mechanical floors, not proof**: they cannot bind the cited source to the specific obligation withdrawn. The zero-loss review in the dual-track gate remains what catches that, and this gate does not pretend to replace it.
|
|
74
|
+
|
|
75
|
+
**This class does not lift the per-owner RED floor.** Seven independent review rounds broke every successive bar built to make an automatic lift safe — a net-byte floor, a normative-line heuristic, a status-letter check, a file-exists check, a substring anchor, and finally a well-formed-but-unbound pointer — and two lanes twice recommended not granting the lift until the pointer can be bound to the obligation actually withdrawn. The residual gap is not machine-checkable in principle: a pointer can resolve to a real heading and quote the deleted text verbatim while the cited source has nothing to do with the claim. So the class does what a machine can do — force an honest label, a real obligation map, and a genuinely pure deletion — and leaves the judgement where it belongs: a withdrawal still needs a `RED-baseline` row, or a named risk owner's waiver through the existing human channel.
|
|
76
|
+
|
|
77
|
+
## Designing a behavioral-evidence measurement
|
|
78
|
+
|
|
79
|
+
A `RED-baseline` row is only as good as the measurement behind it. The failure modes below were each observed while measuring one small skill change; **all of them biased the same way — toward "the change worked" and "the measurement is sound".** That is the tell: when the person designing the measurement is the author of the change being measured, design errors are not random.
|
|
80
|
+
|
|
81
|
+
**The one rule that matters: the grading standard must precede the change.** Not "write the rubric carefully afterwards" — afterwards you already know what the new text says, and the rubric grows into its shape. Either freeze the rubric before editing, or have a party that has not seen the candidate produce it. This repo already applies preregistration to *dispositions* (a preregistered reading rule committed before any run); apply it to the *instrument* too.
|
|
82
|
+
|
|
83
|
+
Everything else is hygiene, but each was observed failing:
|
|
84
|
+
|
|
85
|
+
| Failure | What it looked like | Rule |
|
|
86
|
+
| --- | --- | --- |
|
|
87
|
+
| Ceiling | The unchanged arm already passed 9/10, so no improvement was detectable | The arm you expect to fail must be *able* to fail; verify before comparing |
|
|
88
|
+
| Vocabulary inheritance | Regex written after the new text; the old arm said the same thing in other words and scored 0 (`剥离` vs a regex for `剥掉`) | Grade the **obligation**, paraphrase-tolerant, by a grader blind to which arm produced the answer |
|
|
89
|
+
| Tautology by construction | A "marker must be uniquely carried by this line" rule forced markers that only that line's vocabulary could satisfy — removing the line trivially removed the word | A validity constraint built for one polarity becomes a bias generator in the other: uniqueness is required of the **control**, never of the candidate (other carriers are the redundancy evidence, not an artifact) |
|
|
90
|
+
| Underpowered null | n=3 nulls recorded as findings; one flipped to a large effect at n=10 | Nulls need power and replication; large separations survive small n, nulls do not |
|
|
91
|
+
| Cross-run drift | Identical inputs, same n, unchanged arm scored 2/10 and 6/10 in two runs | Replicate; report the spread, never a point estimate |
|
|
92
|
+
| Transcribed arms | The "old" arm was hand-copied and did not match the real prior text; it inflated the effect | Extract both arms from version control |
|
|
93
|
+
| Slice inflation | Removing one line from a 22-line excerpt overstates its weight versus the 338-line artifact | Use the whole artifact, or report both readings |
|
|
94
|
+
| Post-hoc analysis change | Verdict logic was changed after an unwelcome result | Freeze thresholds and decision order with the rubric; if changed later, publish the full sensitivity grid, never the favourable cell |
|
|
95
|
+
|
|
96
|
+
**Proving a removal is harmless is not the mirror of proving an addition helps.** An addition is evidenced by a difference; a removal is evidenced by an *absence* of difference, which is indistinguishable from an instrument that detects nothing. So a deletion claim needs a **positive control** — remove something known to carry an obligation and show that obligation's satisfaction actually drops. Better still, remove two candidates in the same run so each is the other's control: a double dissociation (removing A drops only A's obligations, removing B only B's) validates the instrument and both verdicts at once.
|
|
97
|
+
|
|
98
|
+
## Instruction-following mechanisms
|
|
99
|
+
|
|
100
|
+
Referenced from `SKILL.md`'s "The mechanism underneath" rule. This section holds the withdrawn predecessor, the sources, and their evidence grade; the entrypoint holds only the operative rules.
|
|
101
|
+
|
|
102
|
+
**Withdrawn (do not cite).** An earlier version said: *prose rules do not reliably co-fire — at any single decision point the agent tends to actively apply roughly the ONE most-salient rule*, with the corollary that *appending a bullet does not ADD compliance — it competes for, and can displace, the same salience slot*. Neither vendor's guidance nor any public benchmark supports a winner-take-all salience slot, and the corollary is contradicted outright: OpenAI's guidance is that a single unequivocal sentence is usually enough to steer the model. The claim rested on one session's self-observation.
|
|
103
|
+
|
|
104
|
+
**Sources and what each actually supports.**
|
|
105
|
+
|
|
106
|
+
| Source | What it gives | Evidence grade |
|
|
107
|
+
| --- | --- | --- |
|
|
108
|
+
| [Anthropic, *Effective context engineering for AI agents*](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents) | finite attention budget; "context rot" reported as gentler in some models but emerging across those tested; smallest-set-of-high-signal-tokens; the right-altitude failure modes | vendor engineering post — no dataset, n, or error bars; version-bound; commercially aligned with context-management tooling |
|
|
109
|
+
| [OpenAI, *GPT-4.1 Prompting Guide*](https://developers.openai.com/cookbook/examples/gpt4-1_prompting_guide) | conflicting instructions tend to resolve to the one nearer the end; instructions at both ends of long context beat either alone; check-conflicts-first; a single clear sentence usually steers | same class; explicitly model-generation-bound ("GPT-4.1 tends to…") |
|
|
110
|
+
| [OpenAI, *GPT-5.1 Prompting Guide*](https://cookbook.openai.com/examples/gpt-5/gpt-5-1_prompting_guide) | check-conflicts-first; a published metaprompt recipe for finding contradictions in your own system prompt | same class |
|
|
111
|
+
| [RECAST](https://arxiv.org/html/2505.19030) | joint satisfaction degrades with constraint count; best model averaged 39.75% all-constraints-satisfied on their benchmark | benchmark paper proposing its own dataset and method — a low baseline flatters the contribution; that number is one hard benchmark's order of magnitude, not a usage failure rate |
|
|
112
|
+
|
|
113
|
+
**Why recency is a hazard, not a rule.** Vendor guidance reports that models *tend to follow* whichever instruction sits later — an observation about behaviour, not a licence to resolve conflicts by position. Two ways position becomes dangerous if read as a rule: a later permissive line beats an earlier stricter one (directly contradicting `Conflict Resolution`, which keeps the stricter data-loss/security/contract guard); and text embedded in **untrusted data** — a diff under review, a retrieved document, tool output — sits later within the same authority level and would win by placement alone, which is prompt injection with extra steps. Treat recency as a bias to design against: put the load-bearing rule where the decision happens, and never let placement confer authority.
|
|
114
|
+
|
|
115
|
+
**How to use them.** Vendor guidance and benchmarks are **hypotheses with good provenance** — they tell you what to test on your own corpus, they do not substitute for testing it. Citing them as settled is the same error as landing an unverified claim; so is overruling one with an underpowered probe.
|
|
@@ -271,6 +271,7 @@ skill 改动后,让 agent 重跑这条 trace,**结构性偏离 = 回归信
|
|
|
271
271
|
- **组合分层**:运行时 = 有序的 profile/bundle 层叠加出来的插件树,每层可用 patch 覆盖任一行配置;能打印出机器实际启动的树来核对。
|
|
272
272
|
- **决策在执行处强制**:schema 省略、prompt 过滤、facade、监听顺序都不是权限边界(已进 `llm-inference-integration` / `product-rd-workflow` 评审清单)。
|
|
273
273
|
- 与之相对,把插件系统的**依赖注入与跨插件解析**做得很重是否值得,业界有争议(一线 harness 作者的公开评价:多数插件互不依赖,复杂 DI 在 90% 场景不带来收益,且跨插件类型仍需另解);本仓不采纳"人人可发的插件生态"作为 skill 分发形态。
|
|
274
|
+
- **可执行落点在 llm-inference-integration,本节只留形态记录**:设计或评审 agent 运行时按该 skill 的 `agent-capability-composition.md`(能力三元、可逆注册、子 agent provider seam)、`agent-tool-dispatch.md`(结果外溢与护栏 wrapper)、`agent-session-persistence.md`(剪枝先于摘要、请求前 checkpoint)、`agent-command-sandbox.md`(policy 单一解析 owner)执行;不得再从本节直接提炼可执行规则——那会与 owner 侧产生双写漂移。
|
|
274
275
|
|
|
275
276
|
不含日志/持久化表述——模型可见内容与事件日志的账目不变量另由安全 owner 参与的独立设计处理。
|
|
276
277
|
|
|
@@ -36,7 +36,9 @@ Two wording rules for whichever form wins:
|
|
|
36
36
|
- **No nuance clauses.** "Don't X unless it matters" reopens the negotiation — appending a single nuance clause to a winning recipe degraded it from consistent to noisy in the same tests. Write a real exception as its own conditional on an observable predicate.
|
|
37
37
|
- **Exemption clauses don't scope.** "This limit doesn't apply to code blocks" still suppresses code blocks; if part of the output must be exempt, restructure the rule so it cannot reach that part.
|
|
38
38
|
|
|
39
|
-
Boundary vs the owning Core Rule's salience mechanism: that rule governs how JOINT requirements are structured (walked enumeration at the firing point; merge-over-append); this table governs the form of a SINGLE rule's text once its landing spot is chosen.
|
|
39
|
+
Boundary vs the owning Core Rule's salience mechanism: that rule governs how JOINT requirements are structured (walked enumeration at the firing point; merge-over-append); this table governs the form of a SINGLE rule's text once its landing spot is chosen.
|
|
40
|
+
|
|
41
|
+
- **When the chosen structure IS a walked pause-point checklist, its design follows checklist practice** (WHO Surgical Safety Checklist implementation manual, 2009 — the operational form of The Checklist Manifesto's killer items): an item earns its slot only by being both critical AND not adequately caught by another mechanism; a pause-point section stays roughly five to nine items, overflow routing to the owning rule text instead of more items; each item binds to one specific, unambiguous action — for a DO-CONFIRM item that action is the confirmation itself, and compound confirmed data is allowed only when ONE concrete confirmation event verifies a coherent tuple (WHO's identity-site-procedure-consent item is one verbal patient-verification event, not a container for unrelated checks); obligations that are independently executable are split into their own items or stay in the owning rule, the item names the specific predicate being confirmed, and a rule-index may serve only as the pointer to that predicate's owning rule, never as a substitute for naming it (the Pre-Completion DO-CONFIRM card's index-not-restatement form works because each indexed canonical rule names what is being confirmed); a new or reshaped gate checklist is exercised on a real task shape before rollout (the design-time operability check's author-dogfood leg); and an item is never deleted merely because it keeps failing or is inconvenient — WHO explicitly discourages removing safety steps because they cannot currently be accomplished; a chronically-failing item routes through the same-class-recurrence keep/delete/narrow/replace decision, not silent removal. Provenance is external — adopt the principle, but transferred evidence does NOT exempt the change from the dual-track gate's mandatory behavioral-evidence row: for a behavior-shaping rule, the `RED-baseline` (recorded incident, or a constructed scenario run against BOTH the unchanged baseline and the changed rule — `dual-track-review-gate.md`) is exactly where the form choice gets tested on YOUR case, with the instrument scaled per `harness-patterns-and-eval.md` §3 (before-after / golden-trace for a single skill edit; the system-wide `skill-behavior-eval` fixture battery is ONLY for always-on layer changes, not routine rule edits). A prohibition chosen for an output-shaping failure, or a recipe replacing a prohibition, is a form choice the source measured as capable of backfiring — it must not land on transferred deltas alone.
|
|
40
42
|
|
|
41
43
|
## Obligation-preservation table (required for any rewrite / retire)
|
|
42
44
|
|