@ccoalm/ccl-skills 0.8.0 → 0.10.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/mobile-quality-release.md +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/manual-invocation-and-prompts.md +6 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/SKILL.md +3 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/multi-tenant-isolation.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/state-machine-task-patterns.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/inference-capacity-operations.md +24 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/llm-client-gateway.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/model-prompt-evaluation.md +4 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/contracts-and-state.md +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/async-lifecycle-and-performance.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +3 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/metrics-conventions.md +8 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/sli-slo-design.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/canary-and-rollout-strategy.md +16 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/promotion-gate-and-review.md +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +12 -12
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/code-review-checklist.md +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/delivery-lifecycle.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/rd-standards-doc-family-checklist.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-system-source-of-truth.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/platform-mobile-patterns.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/tokens-and-components.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-audit.md +8 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/multi-tenant-isolation.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/state-machine-task-patterns.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +12 -15
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/attention-budget-ratchet.md +37 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/description-authoring.md +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +37 -30
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/eval-routing.md +24 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +5 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/rule-consolidation.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +81 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +30 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-contract-anchors.sh +126 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-size-budget.sh +197 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/contract-anchors.tsv +15 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-routing-bank.rb +210 -36
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/extraction_review_gate.sh +3 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/gate_receipt.py +576 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_antipattern_grep_panel.sh +80 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh +99 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +25 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_size_budget.sh +251 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_contract_anchors.sh +196 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_grader_diagnostics.sh +222 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_extraction_review_gate.sh +16 -10
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_frozen_case_sanctity.sh +178 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_frozen_case_sanctity_selfproof.sh +108 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_gate_receipt.sh +431 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_pinned_phrase_mutation_walk.sh +151 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_bank_integrity.sh +86 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh +27 -21
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate_extraction_review_state.py +25 -15
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/classical-test-design-techniques.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc-review-and-prioritization.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/update-lifecycle.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +9 -9
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +5 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/e2e-real-flow-testing.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/integration-contract-testing.md +10 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-topology-and-commands.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +2 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/annotation-driven-revision.md +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/figure-and-table-craft.md +8 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/react-architecture.md +3 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-quality-release.md +37 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-ui-quality.md +10 -1
- package/dist/assets/release.json +127 -67
- package/package.json +1 -1
|
@@ -71,20 +71,52 @@ with open(bank, encoding="utf-8") as fh:
|
|
|
71
71
|
elif rid:
|
|
72
72
|
seen[rid] = lineno
|
|
73
73
|
|
|
74
|
+
# "none" is the runner's negative-control / coverage-gap sentinel
|
|
75
|
+
# (eval-routing-bank.rb): the expected outcome is that no catalog skill
|
|
76
|
+
# claims the utterance. It is an outcome, never a routable target.
|
|
74
77
|
expected = row.get("expected_skill")
|
|
75
|
-
if expected and expected not in installed:
|
|
78
|
+
if expected and expected != "none" and expected not in installed:
|
|
76
79
|
bad(f"line {lineno} ({rid}): expected_skill '{expected}' is not a skill in skills/")
|
|
77
80
|
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
+
# Type-check the ORIGINAL value: `x or []` would first convert an
|
|
82
|
+
# explicitly invalid "" (or 0/false) into a clean empty list, so the
|
|
83
|
+
# deterministic lane would pass a row the Ruby runner rejects.
|
|
84
|
+
must_not = row.get("must_not_route_to")
|
|
85
|
+
if must_not is None:
|
|
86
|
+
must_not = []
|
|
87
|
+
elif not isinstance(must_not, list):
|
|
88
|
+
bad(f"line {lineno} ({rid}): must_not_route_to must be a list, got {must_not!r}")
|
|
81
89
|
must_not = []
|
|
82
90
|
for target in must_not:
|
|
83
|
-
if target
|
|
91
|
+
if target == "none":
|
|
92
|
+
bad(f"line {lineno} ({rid}): must_not_route_to names the sentinel 'none' (an outcome, not a routable target)")
|
|
93
|
+
elif target not in installed:
|
|
84
94
|
bad(f"line {lineno} ({rid}): must_not_route_to names '{target}', not a skill in skills/")
|
|
85
95
|
if target == expected:
|
|
86
96
|
bad(f"line {lineno} ({rid}): '{target}' is both expected_skill and must_not_route_to")
|
|
87
97
|
|
|
98
|
+
acceptable = row.get("acceptable")
|
|
99
|
+
if acceptable is None:
|
|
100
|
+
acceptable = []
|
|
101
|
+
elif not isinstance(acceptable, list):
|
|
102
|
+
bad(f"line {lineno} ({rid}): acceptable must be a list, got {acceptable!r}")
|
|
103
|
+
acceptable = []
|
|
104
|
+
for target in acceptable:
|
|
105
|
+
if target != "none" and target not in installed:
|
|
106
|
+
bad(f"line {lineno} ({rid}): acceptable names '{target}', not a skill in skills/ (nor the 'none' sentinel)")
|
|
107
|
+
if target == expected:
|
|
108
|
+
bad(f"line {lineno} ({rid}): '{target}' is both expected_skill and acceptable (redundant)")
|
|
109
|
+
if target in must_not:
|
|
110
|
+
bad(f"line {lineno} ({rid}): '{target}' is both acceptable and must_not_route_to (contradictory)")
|
|
111
|
+
|
|
112
|
+
# Anti-gaming mirror of the runner's vacuous-row rule: expected +
|
|
113
|
+
# acceptable must leave at least one outcome that would fail the row,
|
|
114
|
+
# or every valid grader selection passes and the fixture fakes green.
|
|
115
|
+
if isinstance(acceptable, list) and expected:
|
|
116
|
+
outcome_space = set(installed) | {"none"}
|
|
117
|
+
if not outcome_space - ({expected} | set(acceptable)):
|
|
118
|
+
bad(f"line {lineno} ({rid}): expected_skill plus acceptable cover every possible outcome (vacuous row)")
|
|
119
|
+
|
|
88
120
|
# A truncated or emptied bank must fail rather than vacuously pass.
|
|
89
121
|
MIN_ROWS = 100
|
|
90
122
|
if rows < MIN_ROWS:
|
|
@@ -203,3 +235,52 @@ print(f"routing_bank_provenance_info: source present on {with_source}/{rows} row
|
|
|
203
235
|
f"(docs/f4-skill-effectiveness-harness.md lists it as required; not enforced here)")
|
|
204
236
|
print(f"routing_bank_integrity_ok rows={rows} (structure only — routing NOT evaluated)")
|
|
205
237
|
PY
|
|
238
|
+
main_rc=$?
|
|
239
|
+
[ "$main_rc" -eq 0 ] || exit "$main_rc"
|
|
240
|
+
|
|
241
|
+
# --- validator self-proof (oracle-can-fail) ----------------------------------
|
|
242
|
+
# A green main run proves nothing about the new field predicates unless the
|
|
243
|
+
# validator is also seen to RED on applied mutants. Run this same script once
|
|
244
|
+
# against a synthetic root carrying one mutant per predicate, and require both
|
|
245
|
+
# the non-zero exit and each predicate's own FAIL message. Guarded against
|
|
246
|
+
# recursion: the inner invocation skips this section.
|
|
247
|
+
if [ -z "${BANK_INTEGRITY_SELFPROOF:-}" ]; then
|
|
248
|
+
SP_ROOT="$(mktemp -d "${TMPDIR:-/tmp}/bank-integrity-selfproof.XXXXXX")"
|
|
249
|
+
trap 'rm -rf "$SP_ROOT"' EXIT
|
|
250
|
+
mkdir -p "$SP_ROOT/skills/testing-strategy" "$SP_ROOT/skills/tighten-doc" "$SP_ROOT/eval"
|
|
251
|
+
printf -- '---\ndescription: stub\n---\n' > "$SP_ROOT/skills/testing-strategy/SKILL.md"
|
|
252
|
+
printf -- '---\ndescription: stub\n---\n' > "$SP_ROOT/skills/tighten-doc/SKILL.md"
|
|
253
|
+
cat > "$SP_ROOT/eval/routing-tasks.jsonl" <<'MUTANTS'
|
|
254
|
+
{"id": "bad-mustnot-none", "utterance": "x", "expected_skill": "testing-strategy", "why_expected": "y", "frozen_at_sha": "root", "must_not_route_to": ["none"]}
|
|
255
|
+
{"id": "bad-acc-restate", "utterance": "x", "expected_skill": "testing-strategy", "why_expected": "y", "frozen_at_sha": "root", "acceptable": ["testing-strategy"]}
|
|
256
|
+
{"id": "bad-acc-unknown", "utterance": "x", "expected_skill": "testing-strategy", "why_expected": "y", "frozen_at_sha": "root", "acceptable": ["not-a-skill"]}
|
|
257
|
+
{"id": "bad-acc-contradict", "utterance": "x", "expected_skill": "testing-strategy", "why_expected": "y", "frozen_at_sha": "root", "acceptable": ["tighten-doc"], "must_not_route_to": ["tighten-doc"]}
|
|
258
|
+
{"id": "bad-acc-empty-string", "utterance": "x", "expected_skill": "testing-strategy", "why_expected": "y", "frozen_at_sha": "root", "acceptable": ""}
|
|
259
|
+
{"id": "bad-mustnot-empty-string", "utterance": "x", "expected_skill": "testing-strategy", "why_expected": "y", "frozen_at_sha": "root", "must_not_route_to": ""}
|
|
260
|
+
{"id": "bad-none-typo", "utterance": "x", "expected_skill": "not-a-skill", "why_expected": "y", "frozen_at_sha": "root"}
|
|
261
|
+
{"id": "bad-universal-pass", "utterance": "x", "expected_skill": "testing-strategy", "why_expected": "y", "frozen_at_sha": "root", "acceptable": ["tighten-doc", "none"]}
|
|
262
|
+
MUTANTS
|
|
263
|
+
sp_out="$(BANK_INTEGRITY_SELFPROOF=1 BANK_ROOT="$SP_ROOT" bash "$0" 2>&1)"
|
|
264
|
+
sp_rc=$?
|
|
265
|
+
if [ "$sp_rc" -eq 0 ]; then
|
|
266
|
+
echo "FAIL: validator self-proof — the mutant bank exited 0 (the oracle cannot fail)" >&2
|
|
267
|
+
exit 1
|
|
268
|
+
fi
|
|
269
|
+
for expect in \
|
|
270
|
+
"bad-mustnot-none): must_not_route_to names the sentinel 'none'" \
|
|
271
|
+
"bad-acc-restate): 'testing-strategy' is both expected_skill and acceptable" \
|
|
272
|
+
"bad-acc-unknown): acceptable names 'not-a-skill'" \
|
|
273
|
+
"bad-acc-contradict): 'tighten-doc' is both acceptable and must_not_route_to" \
|
|
274
|
+
"bad-acc-empty-string): acceptable must be a list, got ''" \
|
|
275
|
+
"bad-mustnot-empty-string): must_not_route_to must be a list, got ''" \
|
|
276
|
+
"bad-none-typo): expected_skill 'not-a-skill' is not a skill in skills/" \
|
|
277
|
+
"bad-universal-pass): expected_skill plus acceptable cover every possible outcome (vacuous row)"; do
|
|
278
|
+
case "$sp_out" in
|
|
279
|
+
*"$expect"*) : ;;
|
|
280
|
+
*) echo "FAIL: validator self-proof — expected mutant message not found: $expect" >&2
|
|
281
|
+
echo "$sp_out" >&2
|
|
282
|
+
exit 1 ;;
|
|
283
|
+
esac
|
|
284
|
+
done
|
|
285
|
+
echo "routing_bank_validator_selfproof_ok: 8 applied mutants red on their own assertions"
|
|
286
|
+
fi
|
|
@@ -140,7 +140,7 @@ def make_fixture(
|
|
|
140
140
|
receipt_mutators=None,
|
|
141
141
|
completion=True,
|
|
142
142
|
completion_mutator=None,
|
|
143
|
-
budget=
|
|
143
|
+
budget=1,
|
|
144
144
|
base_shas=None,
|
|
145
145
|
closeout=None,
|
|
146
146
|
unmatched=0,
|
|
@@ -168,9 +168,9 @@ def make_fixture(
|
|
|
168
168
|
for index, findings in enumerate(receipt_findings, start=1):
|
|
169
169
|
mode = "review" if index == 1 else "challenge"
|
|
170
170
|
status = "findings" if findings else "passed"
|
|
171
|
-
remaining =
|
|
171
|
+
remaining = 2 - index
|
|
172
172
|
if status == "findings":
|
|
173
|
-
state = "post_review_budget" if index ==
|
|
173
|
+
state = "post_review_budget" if index == 2 else "findings_pending"
|
|
174
174
|
else:
|
|
175
175
|
state = "reviewed"
|
|
176
176
|
receipt = {
|
|
@@ -181,7 +181,7 @@ def make_fixture(
|
|
|
181
181
|
"review_chain_tracked": True,
|
|
182
182
|
"review_chain_id": "extraction-chain",
|
|
183
183
|
"autonomous_review_index": index,
|
|
184
|
-
"autonomous_review_budget":
|
|
184
|
+
"autonomous_review_budget": 2,
|
|
185
185
|
"autonomous_reviews_remaining": remaining,
|
|
186
186
|
"autonomous_review_allowed": remaining > 0,
|
|
187
187
|
"challenge_index": 0 if index == 1 else index - 1,
|
|
@@ -221,7 +221,7 @@ def make_fixture(
|
|
|
221
221
|
"review_chain_tracked": True,
|
|
222
222
|
"review_chain_id": final["review_chain_id"],
|
|
223
223
|
"autonomous_review_index": final["autonomous_review_index"],
|
|
224
|
-
"autonomous_review_budget":
|
|
224
|
+
"autonomous_review_budget": 2,
|
|
225
225
|
"autonomous_reviews_remaining": final["autonomous_reviews_remaining"],
|
|
226
226
|
"autonomous_review_allowed": False,
|
|
227
227
|
"challenge_budget": budget,
|
|
@@ -411,9 +411,9 @@ run("float-schema", float_schema["ledger"], 1, "schema_version must be 3")
|
|
|
411
411
|
|
|
412
412
|
float_budget = make_fixture(
|
|
413
413
|
"float-budget",
|
|
414
|
-
receipt_mutators={2: lambda row: row.update(challenge_budget=
|
|
414
|
+
receipt_mutators={2: lambda row: row.update(challenge_budget=1.0)},
|
|
415
415
|
)
|
|
416
|
-
run("float-budget", float_budget["ledger"], 1, "challenge_budget must be
|
|
416
|
+
run("float-budget", float_budget["ledger"], 1, "challenge_budget must be 1")
|
|
417
417
|
|
|
418
418
|
nan_schema = make_fixture("nan-schema")
|
|
419
419
|
nan_schema["ledger"]["schema_version"] = float("nan")
|
|
@@ -493,11 +493,17 @@ scope_mismatch = make_fixture("scope-mismatch", receipt_mutators={2: drift_scope
|
|
|
493
493
|
run("scope-mismatch", scope_mismatch["ledger"], 1, "review scope changed")
|
|
494
494
|
|
|
495
495
|
budget_four = make_fixture("budget-four", budget=4)
|
|
496
|
-
run("budget-four", budget_four["ledger"], 1, "challenge_budget must be
|
|
496
|
+
run("budget-four", budget_four["ledger"], 1, "challenge_budget must be 1")
|
|
497
497
|
|
|
498
498
|
review_only = make_fixture("review-only", receipt_findings=[[]])
|
|
499
499
|
run("review-only", review_only["ledger"], 1, "ready requires at least one tracked challenge")
|
|
500
500
|
|
|
501
|
+
# A caller must not out-run the wrapper budget by supplying extra receipts: a
|
|
502
|
+
# third receipt under budget 1 claims a negative remaining count and would
|
|
503
|
+
# otherwise reach completion validation as a ready-state budget bypass.
|
|
504
|
+
over_budget = make_fixture("over-budget", receipt_findings=[[], [], []])
|
|
505
|
+
run("over-budget", over_budget["ledger"], 1, "exceeds the wrapper budget")
|
|
506
|
+
|
|
501
507
|
missing_complete = make_fixture("missing-complete")
|
|
502
508
|
missing_complete["ledger"]["completion_receipt"] = None
|
|
503
509
|
run("missing-complete", missing_complete["ledger"], 1, "requires a completion receipt")
|
|
@@ -537,7 +543,7 @@ run(
|
|
|
537
543
|
|
|
538
544
|
continuation = make_fixture(
|
|
539
545
|
"continuation",
|
|
540
|
-
receipt_findings=[[finding(1)], [
|
|
546
|
+
receipt_findings=[[finding(1)], [finding(2)]],
|
|
541
547
|
completion=False,
|
|
542
548
|
open_last=True,
|
|
543
549
|
)
|
|
@@ -550,10 +556,10 @@ run(
|
|
|
550
556
|
|
|
551
557
|
continuation_unconsumed_first_drift = make_fixture(
|
|
552
558
|
"continuation-unconsumed-first-drift",
|
|
553
|
-
receipt_findings=[[finding(1)], [
|
|
559
|
+
receipt_findings=[[finding(1)], [finding(2)]],
|
|
554
560
|
completion=False,
|
|
555
561
|
open_last=True,
|
|
556
|
-
base_shas=[A, A,
|
|
562
|
+
base_shas=[A, A, B],
|
|
557
563
|
)
|
|
558
564
|
run(
|
|
559
565
|
"continuation-unconsumed-first-drift",
|
|
@@ -564,7 +570,7 @@ run(
|
|
|
564
570
|
|
|
565
571
|
continuation_candidate = make_fixture(
|
|
566
572
|
"continuation-candidate",
|
|
567
|
-
receipt_findings=[[finding(1)], [
|
|
573
|
+
receipt_findings=[[finding(1)], [finding(2)]],
|
|
568
574
|
completion=False,
|
|
569
575
|
open_last=True,
|
|
570
576
|
)
|
|
@@ -595,8 +601,8 @@ run(
|
|
|
595
601
|
|
|
596
602
|
unknown_state = make_fixture(
|
|
597
603
|
"unknown-state",
|
|
598
|
-
receipt_findings=[[finding(1)], [
|
|
599
|
-
receipt_mutators={
|
|
604
|
+
receipt_findings=[[finding(1)], [finding(2)]],
|
|
605
|
+
receipt_mutators={2: lambda row: row.update(review_state="potato")},
|
|
600
606
|
completion=False,
|
|
601
607
|
open_last=True,
|
|
602
608
|
)
|
|
@@ -605,10 +611,10 @@ run("unknown-state", unknown_state["ledger"], 1, "unknown controller review_stat
|
|
|
605
611
|
|
|
606
612
|
passed_continuation = make_fixture(
|
|
607
613
|
"passed-continuation",
|
|
608
|
-
receipt_findings=[[finding(1)], []
|
|
614
|
+
receipt_findings=[[finding(1)], []],
|
|
609
615
|
completion=False,
|
|
610
616
|
)
|
|
611
|
-
run("passed-continuation", passed_continuation["ledger"], 1, "
|
|
617
|
+
run("passed-continuation", passed_continuation["ledger"], 1, "findings in post_review_budget")
|
|
612
618
|
|
|
613
619
|
bad_finding_hash = make_fixture("bad-finding-hash")
|
|
614
620
|
bad_finding_hash["ledger"]["finding_classes"][0]["occurrences"][0][
|
|
@@ -744,7 +750,7 @@ run(
|
|
|
744
750
|
|
|
745
751
|
historical_open = make_fixture(
|
|
746
752
|
"historical-open",
|
|
747
|
-
receipt_findings=[[finding(1)
|
|
753
|
+
receipt_findings=[[finding(1), finding(2)], []],
|
|
748
754
|
)
|
|
749
755
|
historical_open["ledger"]["finding_classes"][0]["occurrences"][0][
|
|
750
756
|
"disposition"
|
|
@@ -756,7 +762,7 @@ run("historical-open", historical_open["ledger"], 1, "unresolved finding occurre
|
|
|
756
762
|
|
|
757
763
|
historical_open_linked = make_fixture(
|
|
758
764
|
"historical-open-linked",
|
|
759
|
-
receipt_findings=[[finding(1)
|
|
765
|
+
receipt_findings=[[finding(1), finding(2)], []],
|
|
760
766
|
)
|
|
761
767
|
historical_open_linked_occurrences = historical_open_linked["ledger"][
|
|
762
768
|
"finding_classes"
|
|
@@ -825,7 +831,7 @@ run("needs-human", needs_human["ledger"], 1, "unresolved finding occurrence")
|
|
|
825
831
|
|
|
826
832
|
third_occurrence = make_fixture(
|
|
827
833
|
"third-occurrence",
|
|
828
|
-
receipt_findings=[[finding(1)
|
|
834
|
+
receipt_findings=[[finding(1), finding(2)], [finding(3)]],
|
|
829
835
|
completion=False,
|
|
830
836
|
unmatched=1,
|
|
831
837
|
open_last=True,
|
|
@@ -1045,7 +1051,7 @@ for closing_disposition in sorted(CLOSED_DISPOSITIONS):
|
|
|
1045
1051
|
case_name = f"historical-needs-human-{closing_disposition.replace('_', '-')}"
|
|
1046
1052
|
historical_needs_human_linked = make_fixture(
|
|
1047
1053
|
case_name,
|
|
1048
|
-
receipt_findings=[[finding(1)
|
|
1054
|
+
receipt_findings=[[finding(1), finding(2)], []],
|
|
1049
1055
|
)
|
|
1050
1056
|
historical_needs_human_occurrences = historical_needs_human_linked["ledger"][
|
|
1051
1057
|
"finding_classes"
|
|
@@ -1128,7 +1134,7 @@ duplicate_completion_result = invoke(
|
|
|
1128
1134
|
|
|
1129
1135
|
duplicate_sweep = make_fixture(
|
|
1130
1136
|
"duplicate-sweep",
|
|
1131
|
-
receipt_findings=[[finding(1)
|
|
1137
|
+
receipt_findings=[[finding(1), finding(2)], [finding(3)]],
|
|
1132
1138
|
completion=False,
|
|
1133
1139
|
open_last=True,
|
|
1134
1140
|
)
|
|
@@ -30,6 +30,11 @@ TERMINAL_STATES = {
|
|
|
30
30
|
}
|
|
31
31
|
EXTERNAL_REVIEW_STATES = {"reviewed", "findings_pending", "post_review_budget"}
|
|
32
32
|
KNOWN_REVIEW_STATES = EXTERNAL_REVIEW_STATES | {"self_reviewed"}
|
|
33
|
+
# The extraction wrapper fixes the autonomous lane at one review plus one
|
|
34
|
+
# challenge; every numeric bound below derives from these two constants so a
|
|
35
|
+
# future budget change lands in exactly one place.
|
|
36
|
+
WRAPPER_CHALLENGE_BUDGET = 1
|
|
37
|
+
WRAPPER_AUTONOMOUS_ROUNDS = WRAPPER_CHALLENGE_BUDGET + 1
|
|
33
38
|
DISPOSITIONS = {
|
|
34
39
|
"fixed",
|
|
35
40
|
"source_refuted",
|
|
@@ -272,8 +277,8 @@ def validate_scope(receipt: dict[str, Any], label: str) -> str:
|
|
|
272
277
|
]
|
|
273
278
|
if len(normalized_risks) != len(set(normalized_risks)):
|
|
274
279
|
fail(f"{label}.review_scope.risk_tags contains duplicates")
|
|
275
|
-
if scope["challenge_budget"] !=
|
|
276
|
-
fail(f"{label}.review_scope.challenge_budget must be
|
|
280
|
+
if scope["challenge_budget"] != WRAPPER_CHALLENGE_BUDGET or type(scope["challenge_budget"]) is not int:
|
|
281
|
+
fail(f"{label}.review_scope.challenge_budget must be {WRAPPER_CHALLENGE_BUDGET}")
|
|
277
282
|
if (
|
|
278
283
|
scope["wording_only_proof_sha256"] is not None
|
|
279
284
|
or scope["wording_only_scope_sha256"] is not None
|
|
@@ -313,6 +318,11 @@ def validate_controller_receipts(
|
|
|
313
318
|
payload["candidate_sha256"], "ledger.candidate_sha256"
|
|
314
319
|
)
|
|
315
320
|
for expected_index, value in enumerate(refs, start=1):
|
|
321
|
+
# A caller supplying more receipts than the wrapper can mint would
|
|
322
|
+
# otherwise claim a negative remaining count and reach completion
|
|
323
|
+
# validation as a ready-state budget bypass.
|
|
324
|
+
if expected_index > WRAPPER_AUTONOMOUS_ROUNDS:
|
|
325
|
+
fail("controller chain exceeds the wrapper budget")
|
|
316
326
|
ref = exact_object(
|
|
317
327
|
value, {"sequence", "file", "sha256"}, f"controller_receipts[{expected_index - 1}]"
|
|
318
328
|
)
|
|
@@ -352,13 +362,13 @@ def validate_controller_receipts(
|
|
|
352
362
|
receipt.get("challenge_index")
|
|
353
363
|
) is not int:
|
|
354
364
|
fail(f"controller receipt {expected_index} challenge_index is invalid")
|
|
355
|
-
if receipt.get("challenge_budget") !=
|
|
356
|
-
fail(f"controller receipt {expected_index} challenge_budget must be
|
|
357
|
-
if receipt.get("autonomous_review_budget") !=
|
|
365
|
+
if receipt.get("challenge_budget") != WRAPPER_CHALLENGE_BUDGET or type(receipt.get("challenge_budget")) is not int:
|
|
366
|
+
fail(f"controller receipt {expected_index} challenge_budget must be {WRAPPER_CHALLENGE_BUDGET}")
|
|
367
|
+
if receipt.get("autonomous_review_budget") != WRAPPER_AUTONOMOUS_ROUNDS or type(
|
|
358
368
|
receipt.get("autonomous_review_budget")
|
|
359
369
|
) is not int:
|
|
360
|
-
fail(f"controller receipt {expected_index} autonomous_review_budget must be
|
|
361
|
-
expected_remaining =
|
|
370
|
+
fail(f"controller receipt {expected_index} autonomous_review_budget must be {WRAPPER_AUTONOMOUS_ROUNDS}")
|
|
371
|
+
expected_remaining = WRAPPER_AUTONOMOUS_ROUNDS - expected_index
|
|
362
372
|
if receipt.get("autonomous_reviews_remaining") != expected_remaining or type(
|
|
363
373
|
receipt.get("autonomous_reviews_remaining")
|
|
364
374
|
) is not int:
|
|
@@ -415,7 +425,7 @@ def validate_controller_receipts(
|
|
|
415
425
|
fail(f"controller receipt {expected_index} has unknown controller review_state {state}")
|
|
416
426
|
expected_state = (
|
|
417
427
|
"post_review_budget"
|
|
418
|
-
if receipt["status"] == "findings" and expected_index ==
|
|
428
|
+
if receipt["status"] == "findings" and expected_index == WRAPPER_AUTONOMOUS_ROUNDS
|
|
419
429
|
else "findings_pending"
|
|
420
430
|
if receipt["status"] == "findings"
|
|
421
431
|
else "reviewed"
|
|
@@ -469,17 +479,17 @@ def validate_completion_receipt(
|
|
|
469
479
|
fail("completion receipt cannot close a final external receipt with findings")
|
|
470
480
|
if receipt.get("review_chain_tracked") is not True or receipt.get("review_chain_id") != chain_id:
|
|
471
481
|
fail("completion receipt review_chain_id does not match the controller chain")
|
|
472
|
-
if receipt.get("challenge_budget") !=
|
|
473
|
-
fail("completion receipt challenge_budget must be
|
|
474
|
-
if receipt.get("autonomous_review_budget") !=
|
|
482
|
+
if receipt.get("challenge_budget") != WRAPPER_CHALLENGE_BUDGET or type(receipt.get("challenge_budget")) is not int:
|
|
483
|
+
fail(f"completion receipt challenge_budget must be {WRAPPER_CHALLENGE_BUDGET}")
|
|
484
|
+
if receipt.get("autonomous_review_budget") != WRAPPER_AUTONOMOUS_ROUNDS or type(
|
|
475
485
|
receipt.get("autonomous_review_budget")
|
|
476
486
|
) is not int:
|
|
477
|
-
fail("completion receipt autonomous_review_budget must be
|
|
487
|
+
fail(f"completion receipt autonomous_review_budget must be {WRAPPER_AUTONOMOUS_ROUNDS}")
|
|
478
488
|
if receipt.get("autonomous_review_index") != len(receipts) or type(
|
|
479
489
|
receipt.get("autonomous_review_index")
|
|
480
490
|
) is not int:
|
|
481
491
|
fail("completion receipt autonomous_review_index does not match the final round")
|
|
482
|
-
expected_remaining =
|
|
492
|
+
expected_remaining = WRAPPER_AUTONOMOUS_ROUNDS - len(receipts)
|
|
483
493
|
if receipt.get("autonomous_reviews_remaining") != expected_remaining or type(
|
|
484
494
|
receipt.get("autonomous_reviews_remaining")
|
|
485
495
|
) is not int:
|
|
@@ -941,12 +951,12 @@ def validate(payload: dict[str, Any], ledger_dir: Path) -> tuple[str, int, int]:
|
|
|
941
951
|
elif closeout == "continuation_authorization_required":
|
|
942
952
|
final = receipts[-1]
|
|
943
953
|
if (
|
|
944
|
-
len(receipts) !=
|
|
954
|
+
len(receipts) != WRAPPER_AUTONOMOUS_ROUNDS
|
|
945
955
|
or final.get("status") != "findings"
|
|
946
956
|
or final.get("review_state") != "post_review_budget"
|
|
947
957
|
or final.get("human_decision_required") is not True
|
|
948
958
|
):
|
|
949
|
-
fail("continuation_authorization_required requires round
|
|
959
|
+
fail(f"continuation_authorization_required requires final-round (round {WRAPPER_AUTONOMOUS_ROUNDS}) findings in post_review_budget")
|
|
950
960
|
elif closeout == "baseline_race" and not delta:
|
|
951
961
|
fail("baseline_race requires a non-empty unreviewed_delta")
|
|
952
962
|
|
|
@@ -13,7 +13,7 @@
|
|
|
13
13
|
- **James Bach & Michael Bolton**, *Rapid Software Testing* — HTSM / SFDPOT
|
|
14
14
|
- **Elisabeth Hendrickson**, *Explore It!* — test heuristics cheatsheet
|
|
15
15
|
- **James Whittaker**, *Exploratory Software Testing* — tours
|
|
16
|
-
- **Pairwise / Combinatorial**: **Kuhn / Wallace / Gallo 2004 NIST
|
|
16
|
+
- **Pairwise / Combinatorial**: **Kuhn / Wallace / Gallo 2004 NIST 实证**(NIST SP 800-142 Table 1 复现其数据:各被测域 2-way 累计触发 53–97%,多数域 70–97%;NIST 同文提醒 pairwise 仍可能漏掉 10–40% 或更多缺陷,mission-critical 不足恃);工具:Microsoft PICT、NIST ACTS、Hexawise
|
|
17
17
|
- **Hans Buwalda 2004** — soap opera testing
|
|
18
18
|
- **Lisa Crispin & Janet Gregory**, *Agile Testing*
|
|
19
19
|
- **Glenford Myers**, *The Art of Software Testing* (1979) — error guessing 起源
|
|
@@ -61,7 +61,7 @@ SKILL.md 现状 P0 / P1 / P2 是经验判断("blocking / important / nice")
|
|
|
61
61
|
|
|
62
62
|
### 2.1 风险公式(最通用)
|
|
63
63
|
|
|
64
|
-
**Risk = Probability × Impact
|
|
64
|
+
**Risk = Probability × Impact**(出处:ISTQB CTFL v4.0.1 §5.2——风险级别由 likelihood 与 impact 决定,**定量法**为二者相乘,**定性法**用风险矩阵,二者皆合规;ISO/IEC/IEEE 29119-1 采用类似 likelihood/impact 框架,原文付费墙未逐字核)
|
|
65
65
|
|
|
66
66
|
**Probability**(发生概率,1-5 分):
|
|
67
67
|
- 1 = 罕见(新代码 + 简单逻辑 + 测试覆盖好)
|
|
@@ -49,6 +49,8 @@ python test/scripts/gen_report.py \
|
|
|
49
49
|
- md 同步匹配必须以 用例ID 为唯一键;模块名或功能点相同不代表是同一条记录
|
|
50
50
|
- 若 md 先于 Bitable 被修改(如直接编辑文件),需将 md 变更反向同步到 Bitable,再按 update 工作流补信息流转
|
|
51
51
|
- **漂移检查**:定期运行 `python gen_report.py --config test/.report-config.json --diff-md test/cases/all.md` 显示 md 与 Bitable 的 added/removed/changed;废弃记录自动排除(不在 md 里属正常);多人编辑后必跑一次再 commit
|
|
52
|
+
- **版本钉扎**:当仓内测试代码/脚本按某条 TC 实现时,在实现侧记录该 TC 的修订标识——用**定义字段的规范化快照哈希**(或 Bitable 的不可变 record 修订号,若可得);`[姓名 日期]` 不够(同人同日二次修订会撞标识,钉住旧版仍比对相等)。实现前先比对,**回写/提测前再验一次,且回写本身走与认领相同的修订条件更新**(先验后写仍是两步——验证通过与写入之间的并发修改只有条件写能挡)——哈希只有相等性判断:**任何一次不一致都视为分歧,停下、先做源对账**(更新钉扎侧并使差异可评审后再继续);「更新/更旧」的说法仅当平台提供不可变单调修订号时才可用;发现源记录 stale/自相矛盾时**停下先修源记录**,不得按旧版实现后事后补
|
|
53
|
+
- **认领防冲突**:多人/多 agent 并发按 TC 实现时,认领必须是**原子的修订条件更新**(仅当记录修订仍等于读取时的修订才写入——平台支持的 CAS/乐观锁语义)或走**串行化认领协调者**(单写入口)。「读空→写」是 TOCTOU;「写后回读」也不等价——A 回读成功后仍可被 B 顶掉、双方各自都验证通过。两种安全机制都不可得时,**并发认领在该表上不受支持:停下改走人工/单线分配,不得按回读结果继续**。字段已有他人活跃认领时不得覆盖,改为联系认领人或换条目
|
|
52
54
|
- `+record-upsert --record-id` 中的 record_id 是 Bitable 内部 ID,必须先用 `get_record_index()` 从 用例ID 查出,不能直接用 用例ID 代替
|
|
53
55
|
|
|
54
56
|
## update 工作流
|
|
@@ -71,25 +71,25 @@ Use this skill to decide whether a specialized or non-functional test belongs in
|
|
|
71
71
|
- Conditional skips (missing-optional-dependency guards such as module-level `importorskip`, platform/env markers) combined with per-job test selection can leave an entire test file executed in NO CI job while every pipeline stays green: the job that selects the file lacks the optional dependency (the skip fires for the whole module), and the job that has the dependency does not select the file. When a suite mixes conditional skips with job-scoped test selection, the job that owns those tests must carry an executed-count guard — the per-file invariant and the floor fallback live in `references/ci-fixtures-and-flake-control.md`. Any change to job-level selection re-verifies which files each job actually executes (run with skip reporting and read the executed/skipped counts per file). "The tests exist and CI is green" is not evidence they ran anywhere.
|
|
72
72
|
- A new, ported, or mirrored enforcement mechanism (pre-edit hook, permission guard, write-blocking plugin, validator) is not verified by loading, parsing, or config inspection — those prove installation, not enforcement. Require a behavioral matrix before a completion claim — blocked case per deny-condition, allowed/no-collateral cases, the fail-open/swallowed-exception bypass set, and the canonicalization/symlink/worktree edge cases; the matrix cells, the fail-open bypass set, the safe-unavailable-gap disposition for a case that cannot be exercised safely, and the port/mirror parity procedure live in `references/verify-enforcement-mechanisms.md`. Run the matrix only against scratch/synthetic targets (a throwaway checkout/worktree, fixture repo, or dry-run mode) — never a live workspace, real user data, or live credentials. Parity claimed from code reading alone is hypothesis-grade, not evidence. When a test is **ported/mirrored to a sibling stack**, input-fixture fidelity is part of that parity — see `references/test-code-authoring-patterns.md` (跨栈移植:移植对抗输入本身).
|
|
73
73
|
- A "skip CI" / "no runner" instruction does not by itself lower verification rigor, only ceremony. When the blocking CI gate is skipped or unavailable, substitute a same-risk independent check before treating the change as verified — for a tiny/doc/test-only change a local command or `diff --check` is enough; for a change that can break its own gate (it edits the test/tripwire/CI config it is guarded by), an adversarial review/challenge of the diff is what catches the self-break CI would have. Note where a local run is not equivalent to CI (secrets, OS matrix, merge-result pipeline) rather than treating it as full proof. Separately, a project-enforced merge gate (pipeline-must-pass, required review) is not waived by a "skip CI" instruction for convenience: require green status, or an explicit authorized break-glass/override with recorded reason + residual risk — surface the conflict and stop rather than silently bypassing.
|
|
74
|
-
- Browser/E2E smoke must assert visible outcomes, not just click controls. For frontend API pages, verify loading, success, failure, disabled/retry behavior, and absence of dangerous actions where relevant. Capture
|
|
74
|
+
- Browser/E2E smoke must assert visible outcomes, not just click controls. For frontend API pages, verify loading, success, failure, disabled/retry behavior, and absence of dangerous actions where relevant. Capture console errors and failed network requests when tools support it.
|
|
75
75
|
- UI tests and screenshots must prove design quality layers, not only DOM existence. Assert or visually inspect aesthetic hierarchy/density, interaction path, behavioral recovery states, and psychology-critical cues such as disabled reasons, progress certainty, retry safety, confirmation consequences, and return context.
|
|
76
76
|
- For every runtime-visible UI/UX slice, load the canonical sequence in `../product-ui-ux-design/references/delivery-contract.md` and `references/client-runtime-test-matrices.md` §UI/UX Delivery Contract before Phase 0 and after producer/client execution. Testing owns layer selection and sufficiency, binds and cites the complete design/test/producer/client record and candidate-binding sets, confirms every affected client wrote its canonical pre-edit `client_entry` and complete client-record member naming the producer version it exercised, fails closed on a missing/incomplete/mismatched/stale/changed-after-run/unexercised member, and never issues the holistic design verdict.
|
|
77
77
|
- Authentication and account surfaces need an explicit scenario matrix before they can be called complete. Cover identity-input validation across relevant entries, available sign-in methods, registration, account recovery or password reset/change, logout/account switching, sensitive storage/log cleanup, permission or host-authorization denial, and UI/UX acceptance for error copy, disabled reasons, keyboard/safe-area/touch behavior, and visual evidence. If a capability such as recovery, host authorization, real message delivery, or live account verification is absent or external, record it as `product gap`, `blocked`, or `live-only` instead of silently excluding it from the test claim.
|
|
78
|
-
- Test cases come before implementation and broad execution for behavior-changing work. Write a compact test-case register first: scenario, layer, assertion, data/dependency, command, expected current result (`fail`, `pass-existing`, `blocked`, or `gap`), and owner. For bug fixes and user-visible or contract-visible behavior, at least one relevant case must be added or updated and run RED before implementation unless no harness can support it after normal remediation; then record the evidence gap and strongest alternate check. The same RED-first discipline applies to defect records: a reported defect (issue, QA finding) carries the
|
|
79
|
-
- Do not answer "tests are complete" from command output alone.
|
|
78
|
+
- Test cases come before implementation and broad execution for behavior-changing work. Write a compact test-case register first: scenario, layer, assertion, data/dependency, command, expected current result (`fail`, `pass-existing`, `blocked`, `infra-error`, or `gap`), and owner. For bug fixes and user-visible or contract-visible behavior, at least one relevant case must be added or updated and run RED before implementation unless no harness can support it after normal remediation; then record the evidence gap and strongest alternate check. The same RED-first discipline applies to defect records: a reported defect (issue, QA finding) carries the repro command plus the actual failing output, and the fix change references that failing test — a bug "fixed" from its description alone, without a RED reproduction, is unverified.
|
|
79
|
+
- Do not answer "tests are complete" from command output alone. Map each important scenario to a written case or an explicit `blocked`, `live-only`, `product gap`, `infra-error`, or `not applicable` row. (verdict definitions and their mutual exclusivity: `references/ci-fixtures-and-flake-control.md`).
|
|
80
80
|
- For report-only QA, baseline comparison, or "testing only" branches with no product-code changes, failing tests can be the intended deliverable.
|
|
81
81
|
- This exception applies only when a human reviewer, PR owner, or user explicitly states in the current work item, PR description, or current-turn context that the deliverable is test coverage, evidence, or a QA report rather than a product fix; an agent or automated process cannot infer or self-apply this exception from prior-session memory or summarized context.
|
|
82
82
|
- Do not weaken the test or patch product code just to go green.
|
|
83
83
|
- For disputed or high-stakes defects, prefer splitting verification and fix into two deliverables: a test-only verification slice first pins the confirm/deny verdict and root-cause attribution, and the fix is a separate change that references it — keeping the verification verdict uncontaminated by fix intent. Its failing regression test may land skip-marked as the trace only as a bounded state, not an escape: the skip carries the reason, an owner, and the linked fix item, and accepting the fix requires un-skipping it into the blocking regression set (or an explicitly owner-signed quarantine lane) — a RED test that stays skipped after its fix merges is the bypass this rule exists to prevent.
|
|
84
84
|
- The QA report is not complete until pass evidence, red-light evidence, blocked/live-only gaps, and baseline comparison are each present and non-empty or explicitly marked `not applicable`; red-light evidence and baseline comparison cannot both be `not applicable`.
|
|
85
|
-
- Live or production behavior is environment evidence, not a correctness oracle: when live behavior contradicts automated or documented expectations, record the discrepancy as a `live-only gap` with owner
|
|
85
|
+
- Live or production behavior is environment evidence, not a correctness oracle: when live behavior contradicts automated or documented expectations, record the discrepancy as a `live-only gap` with owner; do not resolve it by trusting either side.
|
|
86
86
|
- A QA report with an open live-contradiction gap is `blocked` until the discrepancy is escalated and an owner assigns a resolution path.
|
|
87
|
-
- Generated starter tests are not regression evidence: replace scaffold placeholders
|
|
88
|
-
- Verification warnings are not automatically follow-up work: classify build/bundle
|
|
89
|
-
- Do not open, merge, or describe an MR as ready for a contract-visible change until the test matrix
|
|
87
|
+
- Generated starter tests are not regression evidence: replace scaffold placeholders in the same delivery slice with assertions for the actual app shell, route, state, or user-visible contract.
|
|
88
|
+
- Verification warnings are not automatically follow-up work: classify build/bundle/lint/flaky/deprecation/security/perf warnings from a required gate before reporting success — fix now when caused by the current slice or cheaply local; defer only with reason, residual risk, owner, and follow-up artifact.
|
|
89
|
+
- Do not open, merge, or describe an MR as ready for a contract-visible change until the test matrix is written and executed, or each unavailable layer is explicitly marked unavailable with reason and residual risk. The matrix must include the relevant unit, API/contract, integration, and browser/device/E2E layers; missing layers are release risk, not afterthought.
|
|
90
90
|
- A multi-stack development-standard family is incomplete without a testing standard. The testing standard must define test deliverables, layer policy, harness expectations, CI gates, high-risk coverage, evidence format, and stack handoff rules; stack docs may specialize commands but must not redefine the layer policy.
|
|
91
|
-
- Do not mark a browser/device/E2E layer unavailable just because discovery returns
|
|
92
|
-
- If a browser/device/E2E or host-smoke layer is classified as blocking, unavailable means the delivery is not complete. Use `pre-runtime-test-ready` only when code, lower-layer tests, and build checks are done and a named human/device owner must finish the runtime gate;
|
|
91
|
+
- Do not mark a browser/device/E2E layer unavailable just because discovery returns nothing. First run the normal remediation path: launch the emulator/browser/server/container, wait for readiness, restart the client daemon if appropriate, run the repo setup script, and re-run discovery. Only after that fails may the layer be reported unavailable, with command evidence, residual risk, and next unblock action.
|
|
92
|
+
- If a browser/device/E2E or host-smoke layer is classified as blocking, unavailable means the delivery is not complete. Use `pre-runtime-test-ready` only when code, lower-layer tests, and build checks are done and a named human/device owner must finish the runtime gate; a handoff-only label, not merge-ready/release-ready. Otherwise use `blocked`. Do not describe such work as done, fixed, merge-ready, or release-ready.
|
|
93
93
|
|
|
94
94
|
## Entry Decision: TC Source and Scope
|
|
95
95
|
|
|
@@ -16,7 +16,7 @@ Duplication and dead-code gates run with explicit configuration, not defaults: t
|
|
|
16
16
|
|
|
17
17
|
## Frozen Regression Set And Adversarial Passes
|
|
18
18
|
|
|
19
|
-
Tier the frozen regression set so the gate stays affordable: deterministic frozen cases run in the blocking release gate; cases needing live infra / model calls / real indexes run in the release or pre-ramp gate with an explicit marker, owner, and timeout (do not stuff flaky live cases into the fast gate); human-review-only cases are release evidence, not mislabeled automated tests.
|
|
19
|
+
Tier the frozen regression set so the gate stays affordable: deterministic frozen cases run in the blocking release gate; cases needing live infra / model calls / real indexes run in the release or pre-ramp gate with an explicit marker, owner, and timeout (do not stuff flaky live cases into the fast gate); human-review-only cases are release evidence, not mislabeled automated tests. The tiering criterion is cost and side effects, not speed: a case that consumes paid resources (model/API spend, sandbox creation, render jobs) or mutates external state never belongs in the default always-on lane even when it happens to be fast — the default lane is reserved for cases that are free and side-effect-free to run on every change.
|
|
20
20
|
|
|
21
21
|
The proactive complement — the adversarial pass over code already considered "done": run it as an active defect-discovery step, deliberately hunting coverage blind spots — error-mapping boundaries, concurrency-protection bypass, double-release/double-close paths — instead of waiting for review or production to surface them. Each confirmed gap lands as a failing test first. Candidate blind-spot classes: the risk-matrix failure classes in `scenario-testing.md`, plus the dependency fault-injection and concurrency/cache cases in `integration-contract-testing.md`.
|
|
22
22
|
|
|
@@ -24,6 +24,10 @@ The proactive complement — the adversarial pass over code already considered "
|
|
|
24
24
|
|
|
25
25
|
Use `test-data-and-determinism.md` as the canonical source for fixture shape, anonymization, data builders, golden-file normalization, and deterministic clocks/randomness/ordering.
|
|
26
26
|
|
|
27
|
+
- `infra-error` verdict semantics (the entrypoint's status family). One discriminating predicate decides the verdict — **fault origin**, not symptom: a fault in the **evidence infrastructure** (collector, fixture cache/manifest, state-preparation or controlled-fault harness — anything outside the system under test) = `infra-error`, which is never a pass, never a business fail, and never a silent skip; a fault **in the system under test** (crash, malformed product response, product-path timeout) or an assertion that evaluates to false on collected evidence = business `fail`; an **external prerequisite missing before any attempt** = `blocked`. Exactly one verdict per case; ambiguous origin is resolved by investigation, never by defaulting to whichever verdict looks better; page text, a success toast, or another weaker surface must not substitute for the missing evidence. A required case standing at `infra-error` keeps the aggregate claim incomplete — it counts against merge/release readiness exactly like `blocked`, and a report that excludes `infra-error` cases to present a clean total is a false-green report. Before coding a case, each acceptance criterion names its collector and assertion.
|
|
28
|
+
|
|
29
|
+
External-asset fixtures (media files, documents, large binaries fetched from an external system) form a supply chain that gets pinned end to end: test execution reads only a local read-only cache — never downloads from the external system at run time; a committed manifest pins each asset's identity/hash and CI verifies the manifest plus every blob before the suite runs; cache/manifest verification is part of each affected case's attempted preparation, so a missing or changed cached asset maps to `infra-error` for exactly the cases that need it (a manifest failure aborting before any case attempt marks those cases `infra-error` too, not `blocked` — the infrastructure was configured and failed), never a skip and never a fallback download; seeding/refreshing the cache is a separate offline step on a trusted host, not part of the test run. This composes the network-isolation default and manifest regenerate-and-diff rules in `test-data-and-determinism.md` with the missing-dependency-is-failure rule below into one chain.
|
|
30
|
+
|
|
27
31
|
### Fault-Injection Layers For External-Provider Recovery Paths
|
|
28
32
|
|
|
29
33
|
Recovery behavior against an external provider (a model API, payment/storage backend, streaming dependency) needs its fault permutations proven below the live layer. Layer the fixtures; prove each fault class at the most protocol-real layer that can still script it deterministically:
|
|
@@ -23,7 +23,7 @@ Before adding a new E2E test, create or update the scenario matrix in `scenario-
|
|
|
23
23
|
- Save authenticated state only when the test is not about login.
|
|
24
24
|
- Capture console errors, failed network requests, screenshots/traces/video when useful.
|
|
25
25
|
- Use network interception only to control nondeterminism or assert payloads; do not mock away the contract that the E2E test is meant to prove.
|
|
26
|
-
- **Playwright 1.5x baseline (Microsoft, ongoing 2024-2026)** is the current default-recommendation browser-E2E framework for new web testing. Key features to use deliberately, per `playwright.dev` docs: (a) **Trace Viewer** is the load-bearing debugging surface — every CI failure should produce a `trace.zip` artifact. Set `trace: 'retain-on-failure'` (not `'on-first-retry'`) when CI runs with `retries: 0` for deterministic gating — `'on-first-retry'` produces no trace unless the test actually retries, leaving teams with zero diagnostics on first-failure-then-fix-the-flake debugging cycles. Reviewers open the artifact locally or via `trace.playwright.dev` to step through actions, screenshots, network, and console without re-running the test. (b) **Soft assertions** via `const softExpect = expect.configure({ soft: true })` let one test report multiple failures rather than stopping at the first — useful for state-snapshot assertions where the team wants the whole-page diff in one run, NOT a substitute for the "one test, one behavior" discipline. (c) `toMatchAriaSnapshot()` (Playwright 1.49+) is the structured accessibility-tree assertion — call it on a locator (`expect(page.locator('main')).toMatchAriaSnapshot(...)` or `await page.locator(...).ariaSnapshot()`), preferred over DOM-string snapshots for resilience to non-semantic markup changes. (d) **Component testing
|
|
26
|
+
- **Playwright 1.5x baseline (Microsoft, ongoing 2024-2026)** is the current default-recommendation browser-E2E framework for new web testing (community-adoption evidence, reproducible: `curl -s https://api.npmjs.org/downloads/point/2026-07-31:2026-08-29/<pkg>` for `playwright` / `@playwright/test` / `cypress` returned 339.9M / 216.2M / 30.3M (fixed range, re-verified 2026-08-30), ≈11×; State of JS 2024 testing section, `2024.stateofjs.com/en-US/libraries/testing/`, shows Playwright leading E2E usage/retention — re-run the query before citing as current). Key features to use deliberately, per `playwright.dev` docs: (a) **Trace Viewer** is the load-bearing debugging surface — every CI failure should produce a `trace.zip` artifact. Set `trace: 'retain-on-failure'` (not `'on-first-retry'`) when CI runs with `retries: 0` for deterministic gating — `'on-first-retry'` produces no trace unless the test actually retries, leaving teams with zero diagnostics on first-failure-then-fix-the-flake debugging cycles. Reviewers open the artifact locally or via `trace.playwright.dev` to step through actions, screenshots, network, and console without re-running the test. (b) **Soft assertions** via `const softExpect = expect.configure({ soft: true })` let one test report multiple failures rather than stopping at the first — useful for state-snapshot assertions where the team wants the whole-page diff in one run, NOT a substitute for the "one test, one behavior" discipline. (c) `toMatchAriaSnapshot()` (Playwright 1.49+) is the structured accessibility-tree assertion — call it on a locator (`expect(page.locator('main')).toMatchAriaSnapshot(...)` or `await page.locator(...).ariaSnapshot()`), preferred over DOM-string snapshots for resilience to non-semantic markup changes. (d) **Component testing**: Playwright's current component-testing guide replaces the former `@playwright/experimental-ct-{react,vue}` packages (`playwright.dev/docs/test-components`) — that replacement statement is all the source establishes; draw no stability or package-layout inference from it, and follow the pinned Playwright version's own installation instructions before adding or removing any component-testing package. Vitest browser mode remains a valid component-level alternative when the portfolio already standardizes on Vitest. (e) **Projects** in `playwright.config.ts` define run matrices (browser × device emulation × baseURL × config variant) and produce one merged HTML report. Note: Projects ≠ sharding — sharding is a separate `--shard=k/n` mechanism that splits a single project's tests across multiple workers/machines; teams often combine both (projects for the matrix, sharding for parallelism per project). Pin worker count and shard count for CI determinism, do not let auto-detect choose. Routing the per-stack Playwright config implementation goes to `web-react-dev/references/web-quality-release.md`; this skill owns the test-layer policy.
|
|
27
27
|
|
|
28
28
|
## Runtime QA Sweep
|
|
29
29
|
|
|
@@ -58,7 +58,7 @@ When the deliverable ships as a built or installed artifact — a package `bin`,
|
|
|
58
58
|
- Do not reproduce every unit branch through E2E.
|
|
59
59
|
- Keep E2E flows few, stable, and tied to user/business risk.
|
|
60
60
|
- Prefer one happy path plus high-risk negative paths over many shallow click-throughs.
|
|
61
|
-
- A click-through without assertions is not E2E evidence.
|
|
61
|
+
- A click-through without assertions is not E2E evidence — and weak proxy signals are not business assertions: page loaded, URL changed, non-empty body text, a generic button/canvas/heading visible, or a success toast alone do not prove the business outcome. Anchor the pass condition on objective effects — API response fields, persisted records, balance/count deltas, generated artifact URLs (sufficient alone only when URL issuance is the claimed contract; a generation-success case dereferences the URL and validates artifact status/metadata/content, since a request can issue a valid URL and fail before storing the artifact), or a stable user-visible terminal state (alone only when the visible terminal presentation IS the claimed contract — a rendered "completed" can outrun persistence/billing/artifact creation, so business-outcome cases pair it with the durable effect) — and match assertion strength to what the test title and scenario row claim; a shallow signal is acceptable only when the case explicitly tests just that shallow signal.
|
|
62
62
|
- If a scenario can be proven with a stable API/contract/integration test and only needs one browser smoke for confidence, do not duplicate all permutations in the browser.
|
|
63
63
|
- External-provider fault/recovery permutations (disconnects, malformed streams, rate limits) belong at the protocol-real fault-server and recorded-replay layers (`ci-fixtures-and-flake-control.md`, Fault-Injection Layers); the live credentialed e2e keeps one wiring sanity path, not the fault matrix.
|
|
64
64
|
|
|
@@ -113,6 +113,16 @@ the implementation.
|
|
|
113
113
|
- For cross-RPC typed error envelopes, test a roundtrip: server raises a typed error, the wire-format payload is captured, the client reconstructs a typed error of the same class with the same code/message. Include the unknown-shape path: a wire payload that does not match the canonical envelope returns a transport/unknown error without silent loss of the original cause.
|
|
114
114
|
- For Code-range allocation, test that a service trying to register a code outside its allocated range fails at build/test time, not at runtime.
|
|
115
115
|
|
|
116
|
+
### Cross-Repo Field Change — End-to-End Checklist
|
|
117
|
+
|
|
118
|
+
Adding, renaming, or retyping a field that crosses a repo/service boundary is one end-to-end contract change, not N independent edits. Before calling it covered, walk all five steps (each is a distinct failure site with its own evidence):
|
|
119
|
+
|
|
120
|
+
1. **Producer fallback / bridge** — for an additive field the producer emits a safe default/absent form for consumers that have not upgraded; for a rename/retype a default is NOT enough — keep a compatibility bridge (dual-write old+new representation, or a versioned mapping) until every active consumer is confirmed reading the new form, then remove it in the cleanup stage. Asserted, not assumed.
|
|
121
|
+
2. **Every transport mapper preserves the field explicitly** — do not assume an object spread/copy crosses a mapper or DTO boundary; each mapper in the chain gets an assertion that the field survives it.
|
|
122
|
+
3. **Consumer coverage spans all active consumer variants** — enumerate them from the delivery record's consumer inventory (`../../product-ui-ux-design/references/delivery-contract.md` consumer_inventory for UI variants); testing one variant of a multi-variant consumer is the classic escape.
|
|
123
|
+
4. **One real inbound frame through the mapper, plus one unchanged generic path as control** — the real-frame test proves the new field flows; the untouched-path test proves the change did not perturb everything else (the control catches over-broad mapping edits).
|
|
124
|
+
5. **Paired changes are cross-linked and land in a compatibility-safe order** — not by mutually blocking merges (that deadlocks): consumer tolerance for the field's absence/new form lands first, then the producer emission (its fallback from step 1 keeps not-yet-upgraded consumers safe), then cleanup removes the fallback once all consumers are confirmed upgraded. Each stage gates on the previous stage's **deployed** compatibility state, and the MRs cross-reference per `../../product-rd-workflow/references/cross-repo-coordination.md` for visibility.
|
|
125
|
+
|
|
116
126
|
## Platform Contract / Protobuf / RPC Test Obligations
|
|
117
127
|
|
|
118
128
|
Platform-service-connectivity owns policy and proof mechanics for protobuf-backed HTTP, response envelopes, RPC/base fields, and boundary exposure. `testing-strategy` owns assertion coverage, verdict shape, and CI placement.
|
|
@@ -78,7 +78,7 @@ def test_export_request_returns_signed_url_when_user_has_quota():
|
|
|
78
78
|
|
|
79
79
|
## 3. Test smells(测试异味)
|
|
80
80
|
|
|
81
|
-
**定义**(精选自 Meszaros 2007 及其衍生分类):测试代码的反模式。Meszaros
|
|
81
|
+
**定义**(精选自 Meszaros 2007 及其衍生分类):测试代码的反模式。Meszaros 原书顶层列 15 项 smell,分 code/behavior/project 三类(5/6/4;xunitpatterns.com "All Test Smells" 目录,另有类下变体/别名未计入);下面是日常 review 最常碰到的子集(部分名字 / 阈值是团队启发,非原书字面):
|
|
82
82
|
|
|
83
83
|
| 异味 | 含义 | 后果 |
|
|
84
84
|
|---|---|---|
|
|
@@ -229,7 +229,7 @@ internal_helper_mock.parse.assert_called_once() # 重构改 parse 就挂
|
|
|
229
229
|
- **MC/DC**(Modified Condition/Decision Coverage):每个 boolean 子条件独立影响过决策。**DO-178C 航空 / 医疗 / 汽车 functional safety 才用**;普通业务代码无监管要求时不必上。
|
|
230
230
|
- **Mutation coverage**(见 source-to-case-workflows §C.1):才是真"测得好"的 proxy
|
|
231
231
|
|
|
232
|
-
**用**:CI 设 floor(如 line 60% / branch 50%)防覆盖崩塌;critical-path 模块定专项目标(line 90
|
|
232
|
+
**用**:CI 设 floor(如 line 60% / branch 50%)防覆盖崩塌;critical-path 模块定专项目标(line 90%)。数值出处:60%/90% 对齐 Google Testing Blog "Code Coverage Best Practices"(2020)的 60% acceptable / 75% commendable / 90% exemplary 分档——注意该文同时反对自上而下的强制统一阈值,floor 应按仓现状起步再棘轮;branch 50% 为团队启发值,无外部权威出处。
|
|
233
233
|
|
|
234
234
|
**不用**:
|
|
235
235
|
- **不用 100% 作 KPI** — 强行凑 100% 会产生 lazy assertions(`assert result is not None` 这种)
|
|
@@ -64,7 +64,7 @@ Classify commands before running them:
|
|
|
64
64
|
- E2E/release smoke: browser/API/device real flows in an isolated environment.
|
|
65
65
|
- Long gate: compatibility matrix, migration dry-run, load/replay, visual regression, or full-suite release checks.
|
|
66
66
|
|
|
67
|
-
If the repo separates markers such as `unit`, `integration`, `contract`, `slow`, or `e2e`, preserve that split. Do not move expensive tests into the default PR path unless the local CI contract already expects it.
|
|
67
|
+
If the repo separates markers such as `unit`, `integration`, `contract`, `slow`, or `e2e`, preserve that split. Do not move expensive tests into the default PR path unless the local CI contract already expects it. Expensive includes billable: tests that consume metered external resources — paid AI/model inference, per-call third-party APIs, real payment/checkout flows, cloud sandboxes or device farms — get their own explicit marker/lane, stay out of default PR and scheduled-frequent lanes by default, and run only with a named budget owner and an authorized test account (credential/provisioning discipline per `ci-fixtures-and-flake-control.md`). Payment/checkout tests default to the provider's sandbox/test mode — a budget owner bounds spend but does not make live payment mutation safe; an unavoidable live-money canary needs its own explicit approval, a hard spending cap, and refund/cleanup handling. A case that silently creates paid resources under a generic `e2e` tag is a lane-classification finding.
|
|
68
68
|
|
|
69
69
|
When existing commands use richer labels such as `contract_fake`, `contract_mysql`, `api`, `e2e_smoke`, `failure_mode`, `drill`, `shadow`, `smoke`, or `replay`, preserve the local meaning instead of flattening everything into unit/integration/E2E. Map them to the scenario matrix and CI gate they actually serve.
|
|
70
70
|
|
|
@@ -99,7 +99,7 @@ For a 域卡/执行卡 (a card that sets WHAT a domain must achieve + who owns i
|
|
|
99
99
|
- Required flow: STOP line edits → confirm the corrected core with the user (one short question, don't guess again — this failure class recurs precisely from re-guessing) → re-derive 负责人/红线/里程碑/验收/依赖兜底 from the corrected core as a **draft for review** (not a blind blast-write) → publish on approval.
|
|
100
100
|
- Repeat signal: repeated user "这是什么/什么玩意儿" on the same card = the premise is wrong; escalate to re-derive, do not keep tightening.
|
|
101
101
|
|
|
102
|
-
## 句子层(吸收 Strunk
|
|
102
|
+
## 句子层(吸收 Strunk 与 Google Tech Writing,仅取中文交付文档适用项)
|
|
103
103
|
|
|
104
104
|
> 英文文档:用完整 Strunk 规则(含被本节剔除的语法/标点条),本节只是中文交付子集。
|
|
105
105
|
|
|
@@ -140,6 +140,7 @@ Never destroy collaborative comments. Before editing a collaborative doc, fetch
|
|
|
140
140
|
|
|
141
141
|
## WORKFLOW
|
|
142
142
|
|
|
143
|
+
0. 读者批注:判根因类、全文修同类(`references/annotation-driven-revision.md`)。
|
|
143
144
|
1. Extract the decided-points checklist from the current text.
|
|
144
145
|
2. Apply the DELETE list; keep everything in KEEP.
|
|
145
146
|
3. Rewrite to FORM; confirm every decided point still present.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# 批注驱动修订
|
|
2
|
+
|
|
3
|
+
读者批注/评审意见不是孤立改句请求,而是**阅读断裂的证据**。处理协议:
|
|
4
|
+
|
|
5
|
+
- 逐条判根因类:背景缺失 / 概念未定义 / 逻辑跳跃 / 措辞 / **事实・引用错误**(日期、数字、出处错——此类不走措辞同类扫,改走源核验:对一手源改正并按同源扫其余引用处)。
|
|
6
|
+
- 按根因类**必须全文扫同类位置一起修,不得只改被标记的那一句**——点修复会把同类断裂留给下一位读者复发。**扫描全文、编辑限权**:当授权只覆盖某条批注/某节时,全文扫描产出同类候选清单,但自动编辑只落在授权范围内;范围外的同类位置先报告、经批准再修(不得以「修同类」为名越权改动已定内容)。
|
|
7
|
+
- 改完以首次读者身份通读被改段落(standalone-paste 逐行读,同 closeout 判法)。
|
|
8
|
+
- 批注本体的保全走 SKILL.md 的 COMMENT-SAFE 硬规则(先取真实评论数、定向编辑、改后复核锚点/条数)。
|
|
9
|
+
- 出处:读者差集原则(好文档=读者需要的知识−已有的知识,Google Technical Writing audience 章)——批注正是「差集没算对」的实测信号。
|