@ccoalm/ccl-skills 0.8.0 → 0.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (78) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/mobile-quality-release.md +5 -0
  2. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/manual-invocation-and-prompts.md +6 -0
  3. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +5 -0
  4. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +1 -0
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/SKILL.md +3 -1
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +1 -1
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/multi-tenant-isolation.md +1 -1
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/state-machine-task-patterns.md +2 -0
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/inference-capacity-operations.md +24 -0
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/llm-client-gateway.md +1 -1
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/model-prompt-evaluation.md +4 -1
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/contracts-and-state.md +5 -0
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/async-lifecycle-and-performance.md +1 -0
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +3 -2
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/metrics-conventions.md +8 -1
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/sli-slo-design.md +2 -2
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/canary-and-rollout-strategy.md +16 -2
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/promotion-gate-and-review.md +9 -0
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +12 -12
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/code-review-checklist.md +4 -0
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/delivery-lifecycle.md +1 -1
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/rd-standards-doc-family-checklist.md +1 -0
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-system-source-of-truth.md +2 -0
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/platform-mobile-patterns.md +2 -2
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/tokens-and-components.md +1 -0
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-audit.md +8 -0
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/multi-tenant-isolation.md +1 -1
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/state-machine-task-patterns.md +2 -0
  29. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/SKILL.md +1 -1
  30. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +12 -15
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/attention-budget-ratchet.md +37 -0
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/description-authoring.md +13 -0
  33. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +37 -30
  34. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/eval-routing.md +24 -3
  35. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +5 -5
  36. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/rule-consolidation.md +1 -1
  37. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +81 -0
  38. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +12 -0
  39. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +1 -1
  40. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +30 -0
  41. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-contract-anchors.sh +126 -0
  42. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-size-budget.sh +197 -1
  43. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/contract-anchors.tsv +15 -0
  44. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-routing-bank.rb +210 -36
  45. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/extraction_review_gate.sh +3 -3
  46. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/gate_receipt.py +576 -0
  47. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_antipattern_grep_panel.sh +80 -0
  48. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh +99 -0
  49. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +25 -0
  50. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_size_budget.sh +251 -0
  51. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_contract_anchors.sh +196 -0
  52. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_grader_diagnostics.sh +222 -0
  53. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_extraction_review_gate.sh +16 -10
  54. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_frozen_case_sanctity.sh +178 -0
  55. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_frozen_case_sanctity_selfproof.sh +108 -0
  56. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_gate_receipt.sh +431 -0
  57. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_pinned_phrase_mutation_walk.sh +151 -0
  58. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_bank_integrity.sh +86 -5
  59. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh +27 -21
  60. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate_extraction_review_state.py +25 -15
  61. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/classical-test-design-techniques.md +1 -1
  62. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc-review-and-prioritization.md +1 -1
  63. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/update-lifecycle.md +2 -0
  64. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +9 -9
  65. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +5 -1
  66. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/e2e-real-flow-testing.md +2 -2
  67. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/integration-contract-testing.md +10 -0
  68. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +2 -2
  69. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-topology-and-commands.md +1 -1
  70. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +2 -1
  71. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/annotation-driven-revision.md +9 -0
  72. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/figure-and-table-craft.md +8 -2
  73. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +1 -0
  74. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/react-architecture.md +3 -0
  75. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-quality-release.md +37 -4
  76. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-ui-quality.md +10 -1
  77. package/dist/assets/release.json +127 -67
  78. package/package.json +1 -1
@@ -71,20 +71,52 @@ with open(bank, encoding="utf-8") as fh:
71
71
  elif rid:
72
72
  seen[rid] = lineno
73
73
 
74
+ # "none" is the runner's negative-control / coverage-gap sentinel
75
+ # (eval-routing-bank.rb): the expected outcome is that no catalog skill
76
+ # claims the utterance. It is an outcome, never a routable target.
74
77
  expected = row.get("expected_skill")
75
- if expected and expected not in installed:
78
+ if expected and expected != "none" and expected not in installed:
76
79
  bad(f"line {lineno} ({rid}): expected_skill '{expected}' is not a skill in skills/")
77
80
 
78
- must_not = row.get("must_not_route_to") or []
79
- if not isinstance(must_not, list):
80
- bad(f"line {lineno} ({rid}): must_not_route_to must be a list")
81
+ # Type-check the ORIGINAL value: `x or []` would first convert an
82
+ # explicitly invalid "" (or 0/false) into a clean empty list, so the
83
+ # deterministic lane would pass a row the Ruby runner rejects.
84
+ must_not = row.get("must_not_route_to")
85
+ if must_not is None:
86
+ must_not = []
87
+ elif not isinstance(must_not, list):
88
+ bad(f"line {lineno} ({rid}): must_not_route_to must be a list, got {must_not!r}")
81
89
  must_not = []
82
90
  for target in must_not:
83
- if target not in installed:
91
+ if target == "none":
92
+ bad(f"line {lineno} ({rid}): must_not_route_to names the sentinel 'none' (an outcome, not a routable target)")
93
+ elif target not in installed:
84
94
  bad(f"line {lineno} ({rid}): must_not_route_to names '{target}', not a skill in skills/")
85
95
  if target == expected:
86
96
  bad(f"line {lineno} ({rid}): '{target}' is both expected_skill and must_not_route_to")
87
97
 
98
+ acceptable = row.get("acceptable")
99
+ if acceptable is None:
100
+ acceptable = []
101
+ elif not isinstance(acceptable, list):
102
+ bad(f"line {lineno} ({rid}): acceptable must be a list, got {acceptable!r}")
103
+ acceptable = []
104
+ for target in acceptable:
105
+ if target != "none" and target not in installed:
106
+ bad(f"line {lineno} ({rid}): acceptable names '{target}', not a skill in skills/ (nor the 'none' sentinel)")
107
+ if target == expected:
108
+ bad(f"line {lineno} ({rid}): '{target}' is both expected_skill and acceptable (redundant)")
109
+ if target in must_not:
110
+ bad(f"line {lineno} ({rid}): '{target}' is both acceptable and must_not_route_to (contradictory)")
111
+
112
+ # Anti-gaming mirror of the runner's vacuous-row rule: expected +
113
+ # acceptable must leave at least one outcome that would fail the row,
114
+ # or every valid grader selection passes and the fixture fakes green.
115
+ if isinstance(acceptable, list) and expected:
116
+ outcome_space = set(installed) | {"none"}
117
+ if not outcome_space - ({expected} | set(acceptable)):
118
+ bad(f"line {lineno} ({rid}): expected_skill plus acceptable cover every possible outcome (vacuous row)")
119
+
88
120
  # A truncated or emptied bank must fail rather than vacuously pass.
89
121
  MIN_ROWS = 100
90
122
  if rows < MIN_ROWS:
@@ -203,3 +235,52 @@ print(f"routing_bank_provenance_info: source present on {with_source}/{rows} row
203
235
  f"(docs/f4-skill-effectiveness-harness.md lists it as required; not enforced here)")
204
236
  print(f"routing_bank_integrity_ok rows={rows} (structure only — routing NOT evaluated)")
205
237
  PY
238
+ main_rc=$?
239
+ [ "$main_rc" -eq 0 ] || exit "$main_rc"
240
+
241
+ # --- validator self-proof (oracle-can-fail) ----------------------------------
242
+ # A green main run proves nothing about the new field predicates unless the
243
+ # validator is also seen to RED on applied mutants. Run this same script once
244
+ # against a synthetic root carrying one mutant per predicate, and require both
245
+ # the non-zero exit and each predicate's own FAIL message. Guarded against
246
+ # recursion: the inner invocation skips this section.
247
+ if [ -z "${BANK_INTEGRITY_SELFPROOF:-}" ]; then
248
+ SP_ROOT="$(mktemp -d "${TMPDIR:-/tmp}/bank-integrity-selfproof.XXXXXX")"
249
+ trap 'rm -rf "$SP_ROOT"' EXIT
250
+ mkdir -p "$SP_ROOT/skills/testing-strategy" "$SP_ROOT/skills/tighten-doc" "$SP_ROOT/eval"
251
+ printf -- '---\ndescription: stub\n---\n' > "$SP_ROOT/skills/testing-strategy/SKILL.md"
252
+ printf -- '---\ndescription: stub\n---\n' > "$SP_ROOT/skills/tighten-doc/SKILL.md"
253
+ cat > "$SP_ROOT/eval/routing-tasks.jsonl" <<'MUTANTS'
254
+ {"id": "bad-mustnot-none", "utterance": "x", "expected_skill": "testing-strategy", "why_expected": "y", "frozen_at_sha": "root", "must_not_route_to": ["none"]}
255
+ {"id": "bad-acc-restate", "utterance": "x", "expected_skill": "testing-strategy", "why_expected": "y", "frozen_at_sha": "root", "acceptable": ["testing-strategy"]}
256
+ {"id": "bad-acc-unknown", "utterance": "x", "expected_skill": "testing-strategy", "why_expected": "y", "frozen_at_sha": "root", "acceptable": ["not-a-skill"]}
257
+ {"id": "bad-acc-contradict", "utterance": "x", "expected_skill": "testing-strategy", "why_expected": "y", "frozen_at_sha": "root", "acceptable": ["tighten-doc"], "must_not_route_to": ["tighten-doc"]}
258
+ {"id": "bad-acc-empty-string", "utterance": "x", "expected_skill": "testing-strategy", "why_expected": "y", "frozen_at_sha": "root", "acceptable": ""}
259
+ {"id": "bad-mustnot-empty-string", "utterance": "x", "expected_skill": "testing-strategy", "why_expected": "y", "frozen_at_sha": "root", "must_not_route_to": ""}
260
+ {"id": "bad-none-typo", "utterance": "x", "expected_skill": "not-a-skill", "why_expected": "y", "frozen_at_sha": "root"}
261
+ {"id": "bad-universal-pass", "utterance": "x", "expected_skill": "testing-strategy", "why_expected": "y", "frozen_at_sha": "root", "acceptable": ["tighten-doc", "none"]}
262
+ MUTANTS
263
+ sp_out="$(BANK_INTEGRITY_SELFPROOF=1 BANK_ROOT="$SP_ROOT" bash "$0" 2>&1)"
264
+ sp_rc=$?
265
+ if [ "$sp_rc" -eq 0 ]; then
266
+ echo "FAIL: validator self-proof — the mutant bank exited 0 (the oracle cannot fail)" >&2
267
+ exit 1
268
+ fi
269
+ for expect in \
270
+ "bad-mustnot-none): must_not_route_to names the sentinel 'none'" \
271
+ "bad-acc-restate): 'testing-strategy' is both expected_skill and acceptable" \
272
+ "bad-acc-unknown): acceptable names 'not-a-skill'" \
273
+ "bad-acc-contradict): 'tighten-doc' is both acceptable and must_not_route_to" \
274
+ "bad-acc-empty-string): acceptable must be a list, got ''" \
275
+ "bad-mustnot-empty-string): must_not_route_to must be a list, got ''" \
276
+ "bad-none-typo): expected_skill 'not-a-skill' is not a skill in skills/" \
277
+ "bad-universal-pass): expected_skill plus acceptable cover every possible outcome (vacuous row)"; do
278
+ case "$sp_out" in
279
+ *"$expect"*) : ;;
280
+ *) echo "FAIL: validator self-proof — expected mutant message not found: $expect" >&2
281
+ echo "$sp_out" >&2
282
+ exit 1 ;;
283
+ esac
284
+ done
285
+ echo "routing_bank_validator_selfproof_ok: 8 applied mutants red on their own assertions"
286
+ fi
@@ -140,7 +140,7 @@ def make_fixture(
140
140
  receipt_mutators=None,
141
141
  completion=True,
142
142
  completion_mutator=None,
143
- budget=2,
143
+ budget=1,
144
144
  base_shas=None,
145
145
  closeout=None,
146
146
  unmatched=0,
@@ -168,9 +168,9 @@ def make_fixture(
168
168
  for index, findings in enumerate(receipt_findings, start=1):
169
169
  mode = "review" if index == 1 else "challenge"
170
170
  status = "findings" if findings else "passed"
171
- remaining = 3 - index
171
+ remaining = 2 - index
172
172
  if status == "findings":
173
- state = "post_review_budget" if index == 3 else "findings_pending"
173
+ state = "post_review_budget" if index == 2 else "findings_pending"
174
174
  else:
175
175
  state = "reviewed"
176
176
  receipt = {
@@ -181,7 +181,7 @@ def make_fixture(
181
181
  "review_chain_tracked": True,
182
182
  "review_chain_id": "extraction-chain",
183
183
  "autonomous_review_index": index,
184
- "autonomous_review_budget": 3,
184
+ "autonomous_review_budget": 2,
185
185
  "autonomous_reviews_remaining": remaining,
186
186
  "autonomous_review_allowed": remaining > 0,
187
187
  "challenge_index": 0 if index == 1 else index - 1,
@@ -221,7 +221,7 @@ def make_fixture(
221
221
  "review_chain_tracked": True,
222
222
  "review_chain_id": final["review_chain_id"],
223
223
  "autonomous_review_index": final["autonomous_review_index"],
224
- "autonomous_review_budget": 3,
224
+ "autonomous_review_budget": 2,
225
225
  "autonomous_reviews_remaining": final["autonomous_reviews_remaining"],
226
226
  "autonomous_review_allowed": False,
227
227
  "challenge_budget": budget,
@@ -411,9 +411,9 @@ run("float-schema", float_schema["ledger"], 1, "schema_version must be 3")
411
411
 
412
412
  float_budget = make_fixture(
413
413
  "float-budget",
414
- receipt_mutators={2: lambda row: row.update(challenge_budget=2.0)},
414
+ receipt_mutators={2: lambda row: row.update(challenge_budget=1.0)},
415
415
  )
416
- run("float-budget", float_budget["ledger"], 1, "challenge_budget must be 2")
416
+ run("float-budget", float_budget["ledger"], 1, "challenge_budget must be 1")
417
417
 
418
418
  nan_schema = make_fixture("nan-schema")
419
419
  nan_schema["ledger"]["schema_version"] = float("nan")
@@ -493,11 +493,17 @@ scope_mismatch = make_fixture("scope-mismatch", receipt_mutators={2: drift_scope
493
493
  run("scope-mismatch", scope_mismatch["ledger"], 1, "review scope changed")
494
494
 
495
495
  budget_four = make_fixture("budget-four", budget=4)
496
- run("budget-four", budget_four["ledger"], 1, "challenge_budget must be 2")
496
+ run("budget-four", budget_four["ledger"], 1, "challenge_budget must be 1")
497
497
 
498
498
  review_only = make_fixture("review-only", receipt_findings=[[]])
499
499
  run("review-only", review_only["ledger"], 1, "ready requires at least one tracked challenge")
500
500
 
501
+ # A caller must not out-run the wrapper budget by supplying extra receipts: a
502
+ # third receipt under budget 1 claims a negative remaining count and would
503
+ # otherwise reach completion validation as a ready-state budget bypass.
504
+ over_budget = make_fixture("over-budget", receipt_findings=[[], [], []])
505
+ run("over-budget", over_budget["ledger"], 1, "exceeds the wrapper budget")
506
+
501
507
  missing_complete = make_fixture("missing-complete")
502
508
  missing_complete["ledger"]["completion_receipt"] = None
503
509
  run("missing-complete", missing_complete["ledger"], 1, "requires a completion receipt")
@@ -537,7 +543,7 @@ run(
537
543
 
538
544
  continuation = make_fixture(
539
545
  "continuation",
540
- receipt_findings=[[finding(1)], [], [finding(2)]],
546
+ receipt_findings=[[finding(1)], [finding(2)]],
541
547
  completion=False,
542
548
  open_last=True,
543
549
  )
@@ -550,10 +556,10 @@ run(
550
556
 
551
557
  continuation_unconsumed_first_drift = make_fixture(
552
558
  "continuation-unconsumed-first-drift",
553
- receipt_findings=[[finding(1)], [], [finding(2)]],
559
+ receipt_findings=[[finding(1)], [finding(2)]],
554
560
  completion=False,
555
561
  open_last=True,
556
- base_shas=[A, A, A, B],
562
+ base_shas=[A, A, B],
557
563
  )
558
564
  run(
559
565
  "continuation-unconsumed-first-drift",
@@ -564,7 +570,7 @@ run(
564
570
 
565
571
  continuation_candidate = make_fixture(
566
572
  "continuation-candidate",
567
- receipt_findings=[[finding(1)], [], [finding(2)]],
573
+ receipt_findings=[[finding(1)], [finding(2)]],
568
574
  completion=False,
569
575
  open_last=True,
570
576
  )
@@ -595,8 +601,8 @@ run(
595
601
 
596
602
  unknown_state = make_fixture(
597
603
  "unknown-state",
598
- receipt_findings=[[finding(1)], [], [finding(2)]],
599
- receipt_mutators={3: lambda row: row.update(review_state="potato")},
604
+ receipt_findings=[[finding(1)], [finding(2)]],
605
+ receipt_mutators={2: lambda row: row.update(review_state="potato")},
600
606
  completion=False,
601
607
  open_last=True,
602
608
  )
@@ -605,10 +611,10 @@ run("unknown-state", unknown_state["ledger"], 1, "unknown controller review_stat
605
611
 
606
612
  passed_continuation = make_fixture(
607
613
  "passed-continuation",
608
- receipt_findings=[[finding(1)], [], []],
614
+ receipt_findings=[[finding(1)], []],
609
615
  completion=False,
610
616
  )
611
- run("passed-continuation", passed_continuation["ledger"], 1, "round 3 findings in post_review_budget")
617
+ run("passed-continuation", passed_continuation["ledger"], 1, "findings in post_review_budget")
612
618
 
613
619
  bad_finding_hash = make_fixture("bad-finding-hash")
614
620
  bad_finding_hash["ledger"]["finding_classes"][0]["occurrences"][0][
@@ -744,7 +750,7 @@ run(
744
750
 
745
751
  historical_open = make_fixture(
746
752
  "historical-open",
747
- receipt_findings=[[finding(1)], [finding(2)], []],
753
+ receipt_findings=[[finding(1), finding(2)], []],
748
754
  )
749
755
  historical_open["ledger"]["finding_classes"][0]["occurrences"][0][
750
756
  "disposition"
@@ -756,7 +762,7 @@ run("historical-open", historical_open["ledger"], 1, "unresolved finding occurre
756
762
 
757
763
  historical_open_linked = make_fixture(
758
764
  "historical-open-linked",
759
- receipt_findings=[[finding(1)], [finding(2)], []],
765
+ receipt_findings=[[finding(1), finding(2)], []],
760
766
  )
761
767
  historical_open_linked_occurrences = historical_open_linked["ledger"][
762
768
  "finding_classes"
@@ -825,7 +831,7 @@ run("needs-human", needs_human["ledger"], 1, "unresolved finding occurrence")
825
831
 
826
832
  third_occurrence = make_fixture(
827
833
  "third-occurrence",
828
- receipt_findings=[[finding(1)], [finding(2)], [finding(3)]],
834
+ receipt_findings=[[finding(1), finding(2)], [finding(3)]],
829
835
  completion=False,
830
836
  unmatched=1,
831
837
  open_last=True,
@@ -1045,7 +1051,7 @@ for closing_disposition in sorted(CLOSED_DISPOSITIONS):
1045
1051
  case_name = f"historical-needs-human-{closing_disposition.replace('_', '-')}"
1046
1052
  historical_needs_human_linked = make_fixture(
1047
1053
  case_name,
1048
- receipt_findings=[[finding(1)], [finding(2)], []],
1054
+ receipt_findings=[[finding(1), finding(2)], []],
1049
1055
  )
1050
1056
  historical_needs_human_occurrences = historical_needs_human_linked["ledger"][
1051
1057
  "finding_classes"
@@ -1128,7 +1134,7 @@ duplicate_completion_result = invoke(
1128
1134
 
1129
1135
  duplicate_sweep = make_fixture(
1130
1136
  "duplicate-sweep",
1131
- receipt_findings=[[finding(1)], [finding(2)], [finding(3)]],
1137
+ receipt_findings=[[finding(1), finding(2)], [finding(3)]],
1132
1138
  completion=False,
1133
1139
  open_last=True,
1134
1140
  )
@@ -30,6 +30,11 @@ TERMINAL_STATES = {
30
30
  }
31
31
  EXTERNAL_REVIEW_STATES = {"reviewed", "findings_pending", "post_review_budget"}
32
32
  KNOWN_REVIEW_STATES = EXTERNAL_REVIEW_STATES | {"self_reviewed"}
33
+ # The extraction wrapper fixes the autonomous lane at one review plus one
34
+ # challenge; every numeric bound below derives from these two constants so a
35
+ # future budget change lands in exactly one place.
36
+ WRAPPER_CHALLENGE_BUDGET = 1
37
+ WRAPPER_AUTONOMOUS_ROUNDS = WRAPPER_CHALLENGE_BUDGET + 1
33
38
  DISPOSITIONS = {
34
39
  "fixed",
35
40
  "source_refuted",
@@ -272,8 +277,8 @@ def validate_scope(receipt: dict[str, Any], label: str) -> str:
272
277
  ]
273
278
  if len(normalized_risks) != len(set(normalized_risks)):
274
279
  fail(f"{label}.review_scope.risk_tags contains duplicates")
275
- if scope["challenge_budget"] != 2 or type(scope["challenge_budget"]) is not int:
276
- fail(f"{label}.review_scope.challenge_budget must be 2")
280
+ if scope["challenge_budget"] != WRAPPER_CHALLENGE_BUDGET or type(scope["challenge_budget"]) is not int:
281
+ fail(f"{label}.review_scope.challenge_budget must be {WRAPPER_CHALLENGE_BUDGET}")
277
282
  if (
278
283
  scope["wording_only_proof_sha256"] is not None
279
284
  or scope["wording_only_scope_sha256"] is not None
@@ -313,6 +318,11 @@ def validate_controller_receipts(
313
318
  payload["candidate_sha256"], "ledger.candidate_sha256"
314
319
  )
315
320
  for expected_index, value in enumerate(refs, start=1):
321
+ # A caller supplying more receipts than the wrapper can mint would
322
+ # otherwise claim a negative remaining count and reach completion
323
+ # validation as a ready-state budget bypass.
324
+ if expected_index > WRAPPER_AUTONOMOUS_ROUNDS:
325
+ fail("controller chain exceeds the wrapper budget")
316
326
  ref = exact_object(
317
327
  value, {"sequence", "file", "sha256"}, f"controller_receipts[{expected_index - 1}]"
318
328
  )
@@ -352,13 +362,13 @@ def validate_controller_receipts(
352
362
  receipt.get("challenge_index")
353
363
  ) is not int:
354
364
  fail(f"controller receipt {expected_index} challenge_index is invalid")
355
- if receipt.get("challenge_budget") != 2 or type(receipt.get("challenge_budget")) is not int:
356
- fail(f"controller receipt {expected_index} challenge_budget must be 2")
357
- if receipt.get("autonomous_review_budget") != 3 or type(
365
+ if receipt.get("challenge_budget") != WRAPPER_CHALLENGE_BUDGET or type(receipt.get("challenge_budget")) is not int:
366
+ fail(f"controller receipt {expected_index} challenge_budget must be {WRAPPER_CHALLENGE_BUDGET}")
367
+ if receipt.get("autonomous_review_budget") != WRAPPER_AUTONOMOUS_ROUNDS or type(
358
368
  receipt.get("autonomous_review_budget")
359
369
  ) is not int:
360
- fail(f"controller receipt {expected_index} autonomous_review_budget must be 3")
361
- expected_remaining = 3 - expected_index
370
+ fail(f"controller receipt {expected_index} autonomous_review_budget must be {WRAPPER_AUTONOMOUS_ROUNDS}")
371
+ expected_remaining = WRAPPER_AUTONOMOUS_ROUNDS - expected_index
362
372
  if receipt.get("autonomous_reviews_remaining") != expected_remaining or type(
363
373
  receipt.get("autonomous_reviews_remaining")
364
374
  ) is not int:
@@ -415,7 +425,7 @@ def validate_controller_receipts(
415
425
  fail(f"controller receipt {expected_index} has unknown controller review_state {state}")
416
426
  expected_state = (
417
427
  "post_review_budget"
418
- if receipt["status"] == "findings" and expected_index == 3
428
+ if receipt["status"] == "findings" and expected_index == WRAPPER_AUTONOMOUS_ROUNDS
419
429
  else "findings_pending"
420
430
  if receipt["status"] == "findings"
421
431
  else "reviewed"
@@ -469,17 +479,17 @@ def validate_completion_receipt(
469
479
  fail("completion receipt cannot close a final external receipt with findings")
470
480
  if receipt.get("review_chain_tracked") is not True or receipt.get("review_chain_id") != chain_id:
471
481
  fail("completion receipt review_chain_id does not match the controller chain")
472
- if receipt.get("challenge_budget") != 2 or type(receipt.get("challenge_budget")) is not int:
473
- fail("completion receipt challenge_budget must be 2")
474
- if receipt.get("autonomous_review_budget") != 3 or type(
482
+ if receipt.get("challenge_budget") != WRAPPER_CHALLENGE_BUDGET or type(receipt.get("challenge_budget")) is not int:
483
+ fail(f"completion receipt challenge_budget must be {WRAPPER_CHALLENGE_BUDGET}")
484
+ if receipt.get("autonomous_review_budget") != WRAPPER_AUTONOMOUS_ROUNDS or type(
475
485
  receipt.get("autonomous_review_budget")
476
486
  ) is not int:
477
- fail("completion receipt autonomous_review_budget must be 3")
487
+ fail(f"completion receipt autonomous_review_budget must be {WRAPPER_AUTONOMOUS_ROUNDS}")
478
488
  if receipt.get("autonomous_review_index") != len(receipts) or type(
479
489
  receipt.get("autonomous_review_index")
480
490
  ) is not int:
481
491
  fail("completion receipt autonomous_review_index does not match the final round")
482
- expected_remaining = 3 - len(receipts)
492
+ expected_remaining = WRAPPER_AUTONOMOUS_ROUNDS - len(receipts)
483
493
  if receipt.get("autonomous_reviews_remaining") != expected_remaining or type(
484
494
  receipt.get("autonomous_reviews_remaining")
485
495
  ) is not int:
@@ -941,12 +951,12 @@ def validate(payload: dict[str, Any], ledger_dir: Path) -> tuple[str, int, int]:
941
951
  elif closeout == "continuation_authorization_required":
942
952
  final = receipts[-1]
943
953
  if (
944
- len(receipts) != 3
954
+ len(receipts) != WRAPPER_AUTONOMOUS_ROUNDS
945
955
  or final.get("status") != "findings"
946
956
  or final.get("review_state") != "post_review_budget"
947
957
  or final.get("human_decision_required") is not True
948
958
  ):
949
- fail("continuation_authorization_required requires round 3 findings in post_review_budget")
959
+ fail(f"continuation_authorization_required requires final-round (round {WRAPPER_AUTONOMOUS_ROUNDS}) findings in post_review_budget")
950
960
  elif closeout == "baseline_race" and not delta:
951
961
  fail("baseline_race requires a non-empty unreviewed_delta")
952
962
 
@@ -13,7 +13,7 @@
13
13
  - **James Bach & Michael Bolton**, *Rapid Software Testing* — HTSM / SFDPOT
14
14
  - **Elisabeth Hendrickson**, *Explore It!* — test heuristics cheatsheet
15
15
  - **James Whittaker**, *Exploratory Software Testing* — tours
16
- - **Pairwise / Combinatorial**: **Kuhn / Wallace / Gallo 2004 NIST 实证**(多被测系统 2-way 捕到 50–90% 的缺陷,差异大);工具:Microsoft PICT、NIST ACTS、Hexawise
16
+ - **Pairwise / Combinatorial**: **Kuhn / Wallace / Gallo 2004 NIST 实证**(NIST SP 800-142 Table 1 复现其数据:各被测域 2-way 累计触发 53–97%,多数域 70–97%;NIST 同文提醒 pairwise 仍可能漏掉 10–40% 或更多缺陷,mission-critical 不足恃);工具:Microsoft PICT、NIST ACTS、Hexawise
17
17
  - **Hans Buwalda 2004** — soap opera testing
18
18
  - **Lisa Crispin & Janet Gregory**, *Agile Testing*
19
19
  - **Glenford Myers**, *The Art of Software Testing* (1979) — error guessing 起源
@@ -61,7 +61,7 @@ SKILL.md 现状 P0 / P1 / P2 是经验判断("blocking / important / nice")
61
61
 
62
62
  ### 2.1 风险公式(最通用)
63
63
 
64
- **Risk = Probability × Impact**(业内 30 年共识,ISTQB / ISO 29119 同源)
64
+ **Risk = Probability × Impact**(出处:ISTQB CTFL v4.0.1 §5.2——风险级别由 likelihood 与 impact 决定,**定量法**为二者相乘,**定性法**用风险矩阵,二者皆合规;ISO/IEC/IEEE 29119-1 采用类似 likelihood/impact 框架,原文付费墙未逐字核)
65
65
 
66
66
  **Probability**(发生概率,1-5 分):
67
67
  - 1 = 罕见(新代码 + 简单逻辑 + 测试覆盖好)
@@ -49,6 +49,8 @@ python test/scripts/gen_report.py \
49
49
  - md 同步匹配必须以 用例ID 为唯一键;模块名或功能点相同不代表是同一条记录
50
50
  - 若 md 先于 Bitable 被修改(如直接编辑文件),需将 md 变更反向同步到 Bitable,再按 update 工作流补信息流转
51
51
  - **漂移检查**:定期运行 `python gen_report.py --config test/.report-config.json --diff-md test/cases/all.md` 显示 md 与 Bitable 的 added/removed/changed;废弃记录自动排除(不在 md 里属正常);多人编辑后必跑一次再 commit
52
+ - **版本钉扎**:当仓内测试代码/脚本按某条 TC 实现时,在实现侧记录该 TC 的修订标识——用**定义字段的规范化快照哈希**(或 Bitable 的不可变 record 修订号,若可得);`[姓名 日期]` 不够(同人同日二次修订会撞标识,钉住旧版仍比对相等)。实现前先比对,**回写/提测前再验一次,且回写本身走与认领相同的修订条件更新**(先验后写仍是两步——验证通过与写入之间的并发修改只有条件写能挡)——哈希只有相等性判断:**任何一次不一致都视为分歧,停下、先做源对账**(更新钉扎侧并使差异可评审后再继续);「更新/更旧」的说法仅当平台提供不可变单调修订号时才可用;发现源记录 stale/自相矛盾时**停下先修源记录**,不得按旧版实现后事后补
53
+ - **认领防冲突**:多人/多 agent 并发按 TC 实现时,认领必须是**原子的修订条件更新**(仅当记录修订仍等于读取时的修订才写入——平台支持的 CAS/乐观锁语义)或走**串行化认领协调者**(单写入口)。「读空→写」是 TOCTOU;「写后回读」也不等价——A 回读成功后仍可被 B 顶掉、双方各自都验证通过。两种安全机制都不可得时,**并发认领在该表上不受支持:停下改走人工/单线分配,不得按回读结果继续**。字段已有他人活跃认领时不得覆盖,改为联系认领人或换条目
52
54
  - `+record-upsert --record-id` 中的 record_id 是 Bitable 内部 ID,必须先用 `get_record_index()` 从 用例ID 查出,不能直接用 用例ID 代替
53
55
 
54
56
  ## update 工作流
@@ -71,25 +71,25 @@ Use this skill to decide whether a specialized or non-functional test belongs in
71
71
  - Conditional skips (missing-optional-dependency guards such as module-level `importorskip`, platform/env markers) combined with per-job test selection can leave an entire test file executed in NO CI job while every pipeline stays green: the job that selects the file lacks the optional dependency (the skip fires for the whole module), and the job that has the dependency does not select the file. When a suite mixes conditional skips with job-scoped test selection, the job that owns those tests must carry an executed-count guard — the per-file invariant and the floor fallback live in `references/ci-fixtures-and-flake-control.md`. Any change to job-level selection re-verifies which files each job actually executes (run with skip reporting and read the executed/skipped counts per file). "The tests exist and CI is green" is not evidence they ran anywhere.
72
72
  - A new, ported, or mirrored enforcement mechanism (pre-edit hook, permission guard, write-blocking plugin, validator) is not verified by loading, parsing, or config inspection — those prove installation, not enforcement. Require a behavioral matrix before a completion claim — blocked case per deny-condition, allowed/no-collateral cases, the fail-open/swallowed-exception bypass set, and the canonicalization/symlink/worktree edge cases; the matrix cells, the fail-open bypass set, the safe-unavailable-gap disposition for a case that cannot be exercised safely, and the port/mirror parity procedure live in `references/verify-enforcement-mechanisms.md`. Run the matrix only against scratch/synthetic targets (a throwaway checkout/worktree, fixture repo, or dry-run mode) — never a live workspace, real user data, or live credentials. Parity claimed from code reading alone is hypothesis-grade, not evidence. When a test is **ported/mirrored to a sibling stack**, input-fixture fidelity is part of that parity — see `references/test-code-authoring-patterns.md` (跨栈移植:移植对抗输入本身).
73
73
  - A "skip CI" / "no runner" instruction does not by itself lower verification rigor, only ceremony. When the blocking CI gate is skipped or unavailable, substitute a same-risk independent check before treating the change as verified — for a tiny/doc/test-only change a local command or `diff --check` is enough; for a change that can break its own gate (it edits the test/tripwire/CI config it is guarded by), an adversarial review/challenge of the diff is what catches the self-break CI would have. Note where a local run is not equivalent to CI (secrets, OS matrix, merge-result pipeline) rather than treating it as full proof. Separately, a project-enforced merge gate (pipeline-must-pass, required review) is not waived by a "skip CI" instruction for convenience: require green status, or an explicit authorized break-glass/override with recorded reason + residual risk — surface the conflict and stop rather than silently bypassing.
74
- - Browser/E2E smoke must assert visible outcomes, not just click controls. For frontend API pages, verify loading, success, failure, disabled/retry behavior, and absence of dangerous actions where relevant. Capture or inspect console errors and failed network requests when the browser tool supports it.
74
+ - Browser/E2E smoke must assert visible outcomes, not just click controls. For frontend API pages, verify loading, success, failure, disabled/retry behavior, and absence of dangerous actions where relevant. Capture console errors and failed network requests when tools support it.
75
75
  - UI tests and screenshots must prove design quality layers, not only DOM existence. Assert or visually inspect aesthetic hierarchy/density, interaction path, behavioral recovery states, and psychology-critical cues such as disabled reasons, progress certainty, retry safety, confirmation consequences, and return context.
76
76
  - For every runtime-visible UI/UX slice, load the canonical sequence in `../product-ui-ux-design/references/delivery-contract.md` and `references/client-runtime-test-matrices.md` §UI/UX Delivery Contract before Phase 0 and after producer/client execution. Testing owns layer selection and sufficiency, binds and cites the complete design/test/producer/client record and candidate-binding sets, confirms every affected client wrote its canonical pre-edit `client_entry` and complete client-record member naming the producer version it exercised, fails closed on a missing/incomplete/mismatched/stale/changed-after-run/unexercised member, and never issues the holistic design verdict.
77
77
  - Authentication and account surfaces need an explicit scenario matrix before they can be called complete. Cover identity-input validation across relevant entries, available sign-in methods, registration, account recovery or password reset/change, logout/account switching, sensitive storage/log cleanup, permission or host-authorization denial, and UI/UX acceptance for error copy, disabled reasons, keyboard/safe-area/touch behavior, and visual evidence. If a capability such as recovery, host authorization, real message delivery, or live account verification is absent or external, record it as `product gap`, `blocked`, or `live-only` instead of silently excluding it from the test claim.
78
- - Test cases come before implementation and broad execution for behavior-changing work. Write a compact test-case register first: scenario, layer, assertion, data/dependency, command, expected current result (`fail`, `pass-existing`, `blocked`, or `gap`), and owner. For bug fixes and user-visible or contract-visible behavior, at least one relevant case must be added or updated and run RED before implementation unless no harness can support it after normal remediation; then record the evidence gap and strongest alternate check. The same RED-first discipline applies to defect records: a reported defect (issue, QA finding) carries the reproducible command plus the actual failing output, and the fix change references that failing test — a bug "fixed" from its description alone, without a RED reproduction, is unverified.
79
- - Do not answer "tests are complete" from command output alone. The claim must map each important scenario to a written case or to an explicit `blocked`, `live-only`, `product gap`, or `not applicable` row.
78
+ - Test cases come before implementation and broad execution for behavior-changing work. Write a compact test-case register first: scenario, layer, assertion, data/dependency, command, expected current result (`fail`, `pass-existing`, `blocked`, `infra-error`, or `gap`), and owner. For bug fixes and user-visible or contract-visible behavior, at least one relevant case must be added or updated and run RED before implementation unless no harness can support it after normal remediation; then record the evidence gap and strongest alternate check. The same RED-first discipline applies to defect records: a reported defect (issue, QA finding) carries the repro command plus the actual failing output, and the fix change references that failing test — a bug "fixed" from its description alone, without a RED reproduction, is unverified.
79
+ - Do not answer "tests are complete" from command output alone. Map each important scenario to a written case or an explicit `blocked`, `live-only`, `product gap`, `infra-error`, or `not applicable` row. (verdict definitions and their mutual exclusivity: `references/ci-fixtures-and-flake-control.md`).
80
80
  - For report-only QA, baseline comparison, or "testing only" branches with no product-code changes, failing tests can be the intended deliverable.
81
81
  - This exception applies only when a human reviewer, PR owner, or user explicitly states in the current work item, PR description, or current-turn context that the deliverable is test coverage, evidence, or a QA report rather than a product fix; an agent or automated process cannot infer or self-apply this exception from prior-session memory or summarized context.
82
82
  - Do not weaken the test or patch product code just to go green.
83
83
  - For disputed or high-stakes defects, prefer splitting verification and fix into two deliverables: a test-only verification slice first pins the confirm/deny verdict and root-cause attribution, and the fix is a separate change that references it — keeping the verification verdict uncontaminated by fix intent. Its failing regression test may land skip-marked as the trace only as a bounded state, not an escape: the skip carries the reason, an owner, and the linked fix item, and accepting the fix requires un-skipping it into the blocking regression set (or an explicitly owner-signed quarantine lane) — a RED test that stays skipped after its fix merges is the bypass this rule exists to prevent.
84
84
  - The QA report is not complete until pass evidence, red-light evidence, blocked/live-only gaps, and baseline comparison are each present and non-empty or explicitly marked `not applicable`; red-light evidence and baseline comparison cannot both be `not applicable`.
85
- - Live or production behavior is environment evidence, not a correctness oracle: when live behavior contradicts automated or documented expectations, record the discrepancy as a `live-only gap` with owner and do not resolve it by trusting either side.
85
+ - Live or production behavior is environment evidence, not a correctness oracle: when live behavior contradicts automated or documented expectations, record the discrepancy as a `live-only gap` with owner; do not resolve it by trusting either side.
86
86
  - A QA report with an open live-contradiction gap is `blocked` until the discrepancy is escalated and an owner assigns a resolution path.
87
- - Generated starter tests are not regression evidence: replace scaffold placeholders (e.g. a counter widget test) in the same delivery slice with assertions for the actual app shell, route, state, or user-visible contract.
88
- - Verification warnings are not automatically follow-up work: classify build/bundle-size/lint/flaky/deprecation/security/performance warnings from a required gate before reporting success — fix now when caused by the current slice or cheaply local; defer only with reason, residual risk, owner, and follow-up artifact.
89
- - Do not open, merge, or describe an MR as ready for a contract-visible change until the test matrix has been written and executed, or each unavailable layer is explicitly marked unavailable with reason and residual risk. The matrix must include the relevant unit, API/contract, integration, and browser/device/E2E layers; missing layers are a release risk, not an afterthought.
87
+ - Generated starter tests are not regression evidence: replace scaffold placeholders in the same delivery slice with assertions for the actual app shell, route, state, or user-visible contract.
88
+ - Verification warnings are not automatically follow-up work: classify build/bundle/lint/flaky/deprecation/security/perf warnings from a required gate before reporting success — fix now when caused by the current slice or cheaply local; defer only with reason, residual risk, owner, and follow-up artifact.
89
+ - Do not open, merge, or describe an MR as ready for a contract-visible change until the test matrix is written and executed, or each unavailable layer is explicitly marked unavailable with reason and residual risk. The matrix must include the relevant unit, API/contract, integration, and browser/device/E2E layers; missing layers are release risk, not afterthought.
90
90
  - A multi-stack development-standard family is incomplete without a testing standard. The testing standard must define test deliverables, layer policy, harness expectations, CI gates, high-risk coverage, evidence format, and stack handoff rules; stack docs may specialize commands but must not redefine the layer policy.
91
- - Do not mark a browser/device/E2E layer unavailable just because discovery returns no device, browser, server, or dependency. First run the normal remediation path: launch the emulator/browser/server/container, wait for readiness, restart the client daemon if appropriate, run the repo setup script, and re-run discovery. Only after that fails may the layer be reported unavailable, with command evidence, residual risk, and next unblock action.
92
- - If a browser/device/E2E or host-smoke layer is classified as blocking, unavailable means the delivery is not complete. Use `pre-runtime-test-ready` only when code, lower-layer tests, and build checks are done and a named human/device owner must finish the runtime gate; this is a handoff-only label, not merge-ready or release-ready. Otherwise use `blocked`. Do not describe such work as done, fixed, ready to merge, or ready to release.
91
+ - Do not mark a browser/device/E2E layer unavailable just because discovery returns nothing. First run the normal remediation path: launch the emulator/browser/server/container, wait for readiness, restart the client daemon if appropriate, run the repo setup script, and re-run discovery. Only after that fails may the layer be reported unavailable, with command evidence, residual risk, and next unblock action.
92
+ - If a browser/device/E2E or host-smoke layer is classified as blocking, unavailable means the delivery is not complete. Use `pre-runtime-test-ready` only when code, lower-layer tests, and build checks are done and a named human/device owner must finish the runtime gate; a handoff-only label, not merge-ready/release-ready. Otherwise use `blocked`. Do not describe such work as done, fixed, merge-ready, or release-ready.
93
93
 
94
94
  ## Entry Decision: TC Source and Scope
95
95
 
@@ -16,7 +16,7 @@ Duplication and dead-code gates run with explicit configuration, not defaults: t
16
16
 
17
17
  ## Frozen Regression Set And Adversarial Passes
18
18
 
19
- Tier the frozen regression set so the gate stays affordable: deterministic frozen cases run in the blocking release gate; cases needing live infra / model calls / real indexes run in the release or pre-ramp gate with an explicit marker, owner, and timeout (do not stuff flaky live cases into the fast gate); human-review-only cases are release evidence, not mislabeled automated tests.
19
+ Tier the frozen regression set so the gate stays affordable: deterministic frozen cases run in the blocking release gate; cases needing live infra / model calls / real indexes run in the release or pre-ramp gate with an explicit marker, owner, and timeout (do not stuff flaky live cases into the fast gate); human-review-only cases are release evidence, not mislabeled automated tests. The tiering criterion is cost and side effects, not speed: a case that consumes paid resources (model/API spend, sandbox creation, render jobs) or mutates external state never belongs in the default always-on lane even when it happens to be fast — the default lane is reserved for cases that are free and side-effect-free to run on every change.
20
20
 
21
21
  The proactive complement — the adversarial pass over code already considered "done": run it as an active defect-discovery step, deliberately hunting coverage blind spots — error-mapping boundaries, concurrency-protection bypass, double-release/double-close paths — instead of waiting for review or production to surface them. Each confirmed gap lands as a failing test first. Candidate blind-spot classes: the risk-matrix failure classes in `scenario-testing.md`, plus the dependency fault-injection and concurrency/cache cases in `integration-contract-testing.md`.
22
22
 
@@ -24,6 +24,10 @@ The proactive complement — the adversarial pass over code already considered "
24
24
 
25
25
  Use `test-data-and-determinism.md` as the canonical source for fixture shape, anonymization, data builders, golden-file normalization, and deterministic clocks/randomness/ordering.
26
26
 
27
+ - `infra-error` verdict semantics (the entrypoint's status family). One discriminating predicate decides the verdict — **fault origin**, not symptom: a fault in the **evidence infrastructure** (collector, fixture cache/manifest, state-preparation or controlled-fault harness — anything outside the system under test) = `infra-error`, which is never a pass, never a business fail, and never a silent skip; a fault **in the system under test** (crash, malformed product response, product-path timeout) or an assertion that evaluates to false on collected evidence = business `fail`; an **external prerequisite missing before any attempt** = `blocked`. Exactly one verdict per case; ambiguous origin is resolved by investigation, never by defaulting to whichever verdict looks better; page text, a success toast, or another weaker surface must not substitute for the missing evidence. A required case standing at `infra-error` keeps the aggregate claim incomplete — it counts against merge/release readiness exactly like `blocked`, and a report that excludes `infra-error` cases to present a clean total is a false-green report. Before coding a case, each acceptance criterion names its collector and assertion.
28
+
29
+ External-asset fixtures (media files, documents, large binaries fetched from an external system) form a supply chain that gets pinned end to end: test execution reads only a local read-only cache — never downloads from the external system at run time; a committed manifest pins each asset's identity/hash and CI verifies the manifest plus every blob before the suite runs; cache/manifest verification is part of each affected case's attempted preparation, so a missing or changed cached asset maps to `infra-error` for exactly the cases that need it (a manifest failure aborting before any case attempt marks those cases `infra-error` too, not `blocked` — the infrastructure was configured and failed), never a skip and never a fallback download; seeding/refreshing the cache is a separate offline step on a trusted host, not part of the test run. This composes the network-isolation default and manifest regenerate-and-diff rules in `test-data-and-determinism.md` with the missing-dependency-is-failure rule below into one chain.
30
+
27
31
  ### Fault-Injection Layers For External-Provider Recovery Paths
28
32
 
29
33
  Recovery behavior against an external provider (a model API, payment/storage backend, streaming dependency) needs its fault permutations proven below the live layer. Layer the fixtures; prove each fault class at the most protocol-real layer that can still script it deterministically:
@@ -23,7 +23,7 @@ Before adding a new E2E test, create or update the scenario matrix in `scenario-
23
23
  - Save authenticated state only when the test is not about login.
24
24
  - Capture console errors, failed network requests, screenshots/traces/video when useful.
25
25
  - Use network interception only to control nondeterminism or assert payloads; do not mock away the contract that the E2E test is meant to prove.
26
- - **Playwright 1.5x baseline (Microsoft, ongoing 2024-2026)** is the current default-recommendation browser-E2E framework for new web testing. Key features to use deliberately, per `playwright.dev` docs: (a) **Trace Viewer** is the load-bearing debugging surface — every CI failure should produce a `trace.zip` artifact. Set `trace: 'retain-on-failure'` (not `'on-first-retry'`) when CI runs with `retries: 0` for deterministic gating — `'on-first-retry'` produces no trace unless the test actually retries, leaving teams with zero diagnostics on first-failure-then-fix-the-flake debugging cycles. Reviewers open the artifact locally or via `trace.playwright.dev` to step through actions, screenshots, network, and console without re-running the test. (b) **Soft assertions** via `const softExpect = expect.configure({ soft: true })` let one test report multiple failures rather than stopping at the first — useful for state-snapshot assertions where the team wants the whole-page diff in one run, NOT a substitute for the "one test, one behavior" discipline. (c) `toMatchAriaSnapshot()` (Playwright 1.49+) is the structured accessibility-tree assertion — call it on a locator (`expect(page.locator('main')).toMatchAriaSnapshot(...)` or `await page.locator(...).ariaSnapshot()`), preferred over DOM-string snapshots for resilience to non-semantic markup changes. (d) **Component testing** (`@playwright/experimental-ct-{react,vue,svelte}`) is still labeled experimental — use Vitest browser mode for component-level tests until Playwright component-test API stabilizes. (e) **Projects** in `playwright.config.ts` define run matrices (browser × device emulation × baseURL × config variant) and produce one merged HTML report. Note: Projects ≠ sharding — sharding is a separate `--shard=k/n` mechanism that splits a single project's tests across multiple workers/machines; teams often combine both (projects for the matrix, sharding for parallelism per project). Pin worker count and shard count for CI determinism, do not let auto-detect choose. Routing the per-stack Playwright config implementation goes to `web-react-dev/references/web-quality-release.md`; this skill owns the test-layer policy.
26
+ - **Playwright 1.5x baseline (Microsoft, ongoing 2024-2026)** is the current default-recommendation browser-E2E framework for new web testing (community-adoption evidence, reproducible: `curl -s https://api.npmjs.org/downloads/point/2026-07-31:2026-08-29/<pkg>` for `playwright` / `@playwright/test` / `cypress` returned 339.9M / 216.2M / 30.3M (fixed range, re-verified 2026-08-30), ≈11×; State of JS 2024 testing section, `2024.stateofjs.com/en-US/libraries/testing/`, shows Playwright leading E2E usage/retention — re-run the query before citing as current). Key features to use deliberately, per `playwright.dev` docs: (a) **Trace Viewer** is the load-bearing debugging surface — every CI failure should produce a `trace.zip` artifact. Set `trace: 'retain-on-failure'` (not `'on-first-retry'`) when CI runs with `retries: 0` for deterministic gating — `'on-first-retry'` produces no trace unless the test actually retries, leaving teams with zero diagnostics on first-failure-then-fix-the-flake debugging cycles. Reviewers open the artifact locally or via `trace.playwright.dev` to step through actions, screenshots, network, and console without re-running the test. (b) **Soft assertions** via `const softExpect = expect.configure({ soft: true })` let one test report multiple failures rather than stopping at the first — useful for state-snapshot assertions where the team wants the whole-page diff in one run, NOT a substitute for the "one test, one behavior" discipline. (c) `toMatchAriaSnapshot()` (Playwright 1.49+) is the structured accessibility-tree assertion — call it on a locator (`expect(page.locator('main')).toMatchAriaSnapshot(...)` or `await page.locator(...).ariaSnapshot()`), preferred over DOM-string snapshots for resilience to non-semantic markup changes. (d) **Component testing**: Playwright's current component-testing guide replaces the former `@playwright/experimental-ct-{react,vue}` packages (`playwright.dev/docs/test-components`) — that replacement statement is all the source establishes; draw no stability or package-layout inference from it, and follow the pinned Playwright version's own installation instructions before adding or removing any component-testing package. Vitest browser mode remains a valid component-level alternative when the portfolio already standardizes on Vitest. (e) **Projects** in `playwright.config.ts` define run matrices (browser × device emulation × baseURL × config variant) and produce one merged HTML report. Note: Projects ≠ sharding — sharding is a separate `--shard=k/n` mechanism that splits a single project's tests across multiple workers/machines; teams often combine both (projects for the matrix, sharding for parallelism per project). Pin worker count and shard count for CI determinism, do not let auto-detect choose. Routing the per-stack Playwright config implementation goes to `web-react-dev/references/web-quality-release.md`; this skill owns the test-layer policy.
27
27
 
28
28
  ## Runtime QA Sweep
29
29
 
@@ -58,7 +58,7 @@ When the deliverable ships as a built or installed artifact — a package `bin`,
58
58
  - Do not reproduce every unit branch through E2E.
59
59
  - Keep E2E flows few, stable, and tied to user/business risk.
60
60
  - Prefer one happy path plus high-risk negative paths over many shallow click-throughs.
61
- - A click-through without assertions is not E2E evidence.
61
+ - A click-through without assertions is not E2E evidence — and weak proxy signals are not business assertions: page loaded, URL changed, non-empty body text, a generic button/canvas/heading visible, or a success toast alone do not prove the business outcome. Anchor the pass condition on objective effects — API response fields, persisted records, balance/count deltas, generated artifact URLs (sufficient alone only when URL issuance is the claimed contract; a generation-success case dereferences the URL and validates artifact status/metadata/content, since a request can issue a valid URL and fail before storing the artifact), or a stable user-visible terminal state (alone only when the visible terminal presentation IS the claimed contract — a rendered "completed" can outrun persistence/billing/artifact creation, so business-outcome cases pair it with the durable effect) — and match assertion strength to what the test title and scenario row claim; a shallow signal is acceptable only when the case explicitly tests just that shallow signal.
62
62
  - If a scenario can be proven with a stable API/contract/integration test and only needs one browser smoke for confidence, do not duplicate all permutations in the browser.
63
63
  - External-provider fault/recovery permutations (disconnects, malformed streams, rate limits) belong at the protocol-real fault-server and recorded-replay layers (`ci-fixtures-and-flake-control.md`, Fault-Injection Layers); the live credentialed e2e keeps one wiring sanity path, not the fault matrix.
64
64
 
@@ -113,6 +113,16 @@ the implementation.
113
113
  - For cross-RPC typed error envelopes, test a roundtrip: server raises a typed error, the wire-format payload is captured, the client reconstructs a typed error of the same class with the same code/message. Include the unknown-shape path: a wire payload that does not match the canonical envelope returns a transport/unknown error without silent loss of the original cause.
114
114
  - For Code-range allocation, test that a service trying to register a code outside its allocated range fails at build/test time, not at runtime.
115
115
 
116
+ ### Cross-Repo Field Change — End-to-End Checklist
117
+
118
+ Adding, renaming, or retyping a field that crosses a repo/service boundary is one end-to-end contract change, not N independent edits. Before calling it covered, walk all five steps (each is a distinct failure site with its own evidence):
119
+
120
+ 1. **Producer fallback / bridge** — for an additive field the producer emits a safe default/absent form for consumers that have not upgraded; for a rename/retype a default is NOT enough — keep a compatibility bridge (dual-write old+new representation, or a versioned mapping) until every active consumer is confirmed reading the new form, then remove it in the cleanup stage. Asserted, not assumed.
121
+ 2. **Every transport mapper preserves the field explicitly** — do not assume an object spread/copy crosses a mapper or DTO boundary; each mapper in the chain gets an assertion that the field survives it.
122
+ 3. **Consumer coverage spans all active consumer variants** — enumerate them from the delivery record's consumer inventory (`../../product-ui-ux-design/references/delivery-contract.md` consumer_inventory for UI variants); testing one variant of a multi-variant consumer is the classic escape.
123
+ 4. **One real inbound frame through the mapper, plus one unchanged generic path as control** — the real-frame test proves the new field flows; the untouched-path test proves the change did not perturb everything else (the control catches over-broad mapping edits).
124
+ 5. **Paired changes are cross-linked and land in a compatibility-safe order** — not by mutually blocking merges (that deadlocks): consumer tolerance for the field's absence/new form lands first, then the producer emission (its fallback from step 1 keeps not-yet-upgraded consumers safe), then cleanup removes the fallback once all consumers are confirmed upgraded. Each stage gates on the previous stage's **deployed** compatibility state, and the MRs cross-reference per `../../product-rd-workflow/references/cross-repo-coordination.md` for visibility.
125
+
116
126
  ## Platform Contract / Protobuf / RPC Test Obligations
117
127
 
118
128
  Platform-service-connectivity owns policy and proof mechanics for protobuf-backed HTTP, response envelopes, RPC/base fields, and boundary exposure. `testing-strategy` owns assertion coverage, verdict shape, and CI placement.
@@ -78,7 +78,7 @@ def test_export_request_returns_signed_url_when_user_has_quota():
78
78
 
79
79
  ## 3. Test smells(测试异味)
80
80
 
81
- **定义**(精选自 Meszaros 2007 及其衍生分类):测试代码的反模式。Meszaros 原书列约 18 项分 code/behavior/project 三类;下面是日常 review 最常碰到的子集(部分名字 / 阈值是团队启发,非原书字面):
81
+ **定义**(精选自 Meszaros 2007 及其衍生分类):测试代码的反模式。Meszaros 原书顶层列 15 项 smell,分 code/behavior/project 三类(5/6/4;xunitpatterns.com "All Test Smells" 目录,另有类下变体/别名未计入);下面是日常 review 最常碰到的子集(部分名字 / 阈值是团队启发,非原书字面):
82
82
 
83
83
  | 异味 | 含义 | 后果 |
84
84
  |---|---|---|
@@ -229,7 +229,7 @@ internal_helper_mock.parse.assert_called_once() # 重构改 parse 就挂
229
229
  - **MC/DC**(Modified Condition/Decision Coverage):每个 boolean 子条件独立影响过决策。**DO-178C 航空 / 医疗 / 汽车 functional safety 才用**;普通业务代码无监管要求时不必上。
230
230
  - **Mutation coverage**(见 source-to-case-workflows §C.1):才是真"测得好"的 proxy
231
231
 
232
- **用**:CI 设 floor(如 line 60% / branch 50%)防覆盖崩塌;critical-path 模块定专项目标(line 90%)。
232
+ **用**:CI 设 floor(如 line 60% / branch 50%)防覆盖崩塌;critical-path 模块定专项目标(line 90%)。数值出处:60%/90% 对齐 Google Testing Blog "Code Coverage Best Practices"(2020)的 60% acceptable / 75% commendable / 90% exemplary 分档——注意该文同时反对自上而下的强制统一阈值,floor 应按仓现状起步再棘轮;branch 50% 为团队启发值,无外部权威出处。
233
233
 
234
234
  **不用**:
235
235
  - **不用 100% 作 KPI** — 强行凑 100% 会产生 lazy assertions(`assert result is not None` 这种)
@@ -64,7 +64,7 @@ Classify commands before running them:
64
64
  - E2E/release smoke: browser/API/device real flows in an isolated environment.
65
65
  - Long gate: compatibility matrix, migration dry-run, load/replay, visual regression, or full-suite release checks.
66
66
 
67
- If the repo separates markers such as `unit`, `integration`, `contract`, `slow`, or `e2e`, preserve that split. Do not move expensive tests into the default PR path unless the local CI contract already expects it.
67
+ If the repo separates markers such as `unit`, `integration`, `contract`, `slow`, or `e2e`, preserve that split. Do not move expensive tests into the default PR path unless the local CI contract already expects it. Expensive includes billable: tests that consume metered external resources — paid AI/model inference, per-call third-party APIs, real payment/checkout flows, cloud sandboxes or device farms — get their own explicit marker/lane, stay out of default PR and scheduled-frequent lanes by default, and run only with a named budget owner and an authorized test account (credential/provisioning discipline per `ci-fixtures-and-flake-control.md`). Payment/checkout tests default to the provider's sandbox/test mode — a budget owner bounds spend but does not make live payment mutation safe; an unavoidable live-money canary needs its own explicit approval, a hard spending cap, and refund/cleanup handling. A case that silently creates paid resources under a generic `e2e` tag is a lane-classification finding.
68
68
 
69
69
  When existing commands use richer labels such as `contract_fake`, `contract_mysql`, `api`, `e2e_smoke`, `failure_mode`, `drill`, `shadow`, `smoke`, or `replay`, preserve the local meaning instead of flattening everything into unit/integration/E2E. Map them to the scenario matrix and CI gate they actually serve.
70
70
 
@@ -99,7 +99,7 @@ For a 域卡/执行卡 (a card that sets WHAT a domain must achieve + who owns i
99
99
  - Required flow: STOP line edits → confirm the corrected core with the user (one short question, don't guess again — this failure class recurs precisely from re-guessing) → re-derive 负责人/红线/里程碑/验收/依赖兜底 from the corrected core as a **draft for review** (not a blind blast-write) → publish on approval.
100
100
  - Repeat signal: repeated user "这是什么/什么玩意儿" on the same card = the premise is wrong; escalate to re-derive, do not keep tightening.
101
101
 
102
- ## 句子层(吸收 Strunk《风格的要素》与 Google Technical Writing 课程,仅取适合中文交付文档的;英文语法/标点规则不适用,已剔除)
102
+ ## 句子层(吸收 Strunk 与 Google Tech Writing,仅取中文交付文档适用项)
103
103
 
104
104
  > 英文文档:用完整 Strunk 规则(含被本节剔除的语法/标点条),本节只是中文交付子集。
105
105
 
@@ -140,6 +140,7 @@ Never destroy collaborative comments. Before editing a collaborative doc, fetch
140
140
 
141
141
  ## WORKFLOW
142
142
 
143
+ 0. 读者批注:判根因类、全文修同类(`references/annotation-driven-revision.md`)。
143
144
  1. Extract the decided-points checklist from the current text.
144
145
  2. Apply the DELETE list; keep everything in KEEP.
145
146
  3. Rewrite to FORM; confirm every decided point still present.
@@ -0,0 +1,9 @@
1
+ # 批注驱动修订
2
+
3
+ 读者批注/评审意见不是孤立改句请求,而是**阅读断裂的证据**。处理协议:
4
+
5
+ - 逐条判根因类:背景缺失 / 概念未定义 / 逻辑跳跃 / 措辞 / **事实・引用错误**(日期、数字、出处错——此类不走措辞同类扫,改走源核验:对一手源改正并按同源扫其余引用处)。
6
+ - 按根因类**必须全文扫同类位置一起修,不得只改被标记的那一句**——点修复会把同类断裂留给下一位读者复发。**扫描全文、编辑限权**:当授权只覆盖某条批注/某节时,全文扫描产出同类候选清单,但自动编辑只落在授权范围内;范围外的同类位置先报告、经批准再修(不得以「修同类」为名越权改动已定内容)。
7
+ - 改完以首次读者身份通读被改段落(standalone-paste 逐行读,同 closeout 判法)。
8
+ - 批注本体的保全走 SKILL.md 的 COMMENT-SAFE 硬规则(先取真实评论数、定向编辑、改后复核锚点/条数)。
9
+ - 出处:读者差集原则(好文档=读者需要的知识−已有的知识,Google Technical Writing audience 章)——批注正是「差集没算对」的实测信号。