@ccoalm/ccl-skills 0.15.1 → 0.15.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (50) hide show
  1. package/README.md +3 -1
  2. package/dist/assets/marketplace/plugins/ccl-skills/agent-context/session-start.md +1 -1
  3. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/SKILL.md +2 -2
  4. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +6 -5
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/development-completion.md +26 -0
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +37 -13
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/AGENTS.md +5 -2
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/codex_review.sh +77 -5
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/kimi_packet_mcp.py +98 -4
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_cli_review.py +48 -1
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +230 -16
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_cli_review_wrappers.sh +165 -11
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_kimi_packet_mcp.py +143 -0
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_compat.py +572 -0
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +3 -1
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/SKILL.md +1 -1
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +2 -0
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/SKILL.md +1 -1
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-agent-delegation/SKILL.md +1 -1
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/SKILL.md +2 -0
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +3 -1
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/SKILL.md +3 -1
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/SKILL.md +2 -0
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +7 -7
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-review-gate-mechanics.md +1 -1
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/pre-final-continuation-gate.md +20 -11
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/refactoring-discipline.md +7 -1
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/SKILL.md +1 -1
  29. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +2 -2
  30. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +17 -17
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +4 -4
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/resume-paused-delivery.md +3 -3
  33. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +11 -0
  34. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh +83 -48
  35. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh +80 -2
  36. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_controlled_escalation_pins.sh +3 -2
  37. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh +190 -2
  38. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate_extraction_review_state.py +106 -4
  39. package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/SKILL.md +1 -1
  40. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +3 -1
  41. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +1 -1
  42. package/dist/assets/release.json +54 -49
  43. package/dist/codex-host.d.ts +1 -1
  44. package/dist/codex-host.js +39 -15
  45. package/dist/host-probe.d.ts +16 -0
  46. package/dist/host-probe.js +29 -5
  47. package/dist/operations.js +34 -8
  48. package/dist/unified.d.ts +1 -1
  49. package/dist/unified.js +11 -4
  50. package/package.json +1 -1
@@ -4,7 +4,11 @@
4
4
  from __future__ import annotations
5
5
 
6
6
  import argparse
7
+ import contextlib
8
+ import copy
9
+ import hashlib
7
10
  import importlib.util
11
+ import io
8
12
  import json
9
13
  import os
10
14
  from pathlib import Path
@@ -12,12 +16,14 @@ import subprocess
12
16
  import sys
13
17
  import tempfile
14
18
  import unittest
19
+ from unittest import mock
15
20
 
16
21
 
17
22
  SCRIPT_DIR = Path(__file__).resolve().parent
18
23
  if str(SCRIPT_DIR) not in sys.path:
19
24
  sys.path.insert(0, str(SCRIPT_DIR))
20
25
  import review_gate
26
+ import kimi_packet_mcp
21
27
 
22
28
  SPEC = importlib.util.spec_from_file_location(
23
29
  "parse_cli_review", SCRIPT_DIR / "parse_cli_review.py"
@@ -505,5 +511,571 @@ class ReviewClientCompatibilityTest(unittest.TestCase):
505
511
  self.assertFalse(eligibility)
506
512
 
507
513
 
514
+ class CodexPacketAuditTest(unittest.TestCase):
515
+ def setUp(self) -> None:
516
+ temporary = tempfile.TemporaryDirectory()
517
+ self.addCleanup(temporary.cleanup)
518
+ self.packet = Path(temporary.name) / "packet.txt"
519
+ self.packet.write_text("candidate\n", encoding="utf-8")
520
+ result = Path(temporary.name) / "result.txt"
521
+ result.write_text("NO_BLOCKING_FINDINGS\n", encoding="utf-8")
522
+ self.digest = hashlib.sha256(self.packet.read_bytes()).hexdigest()
523
+ self.args = argparse.Namespace(
524
+ client="codex",
525
+ mode="review",
526
+ reviewer_family="openai",
527
+ provider="codex-cli",
528
+ model="",
529
+ result_file=str(result),
530
+ packet=str(self.packet),
531
+ packet_sha256=self.digest,
532
+ )
533
+
534
+ @staticmethod
535
+ def completed_events(*items: dict) -> list[dict]:
536
+ return [
537
+ {"type": "thread.started", "thread_id": "synthetic-review"},
538
+ {"type": "turn.started"},
539
+ *items,
540
+ {
541
+ "type": "item.completed",
542
+ "item": {
543
+ "id": "verdict",
544
+ "type": "agent_message",
545
+ "text": "NO_BLOCKING_FINDINGS",
546
+ },
547
+ },
548
+ {"type": "turn.completed"},
549
+ ]
550
+
551
+ @staticmethod
552
+ def read_event() -> dict:
553
+ return {
554
+ "type": "item.completed",
555
+ "item": {
556
+ "id": "read-1",
557
+ "type": "mcp_tool_call",
558
+ "server": "code_review_packet",
559
+ "tool": "read_packet",
560
+ "arguments": {"byte_offset": 0, "max_bytes": 46_000},
561
+ "status": "completed",
562
+ "result": {
563
+ "content": [
564
+ {"type": "text", "text": "PACKET_CHUNK 0:10/10\ncandidate\n"}
565
+ ],
566
+ "structured_content": None,
567
+ },
568
+ "error": None,
569
+ },
570
+ }
571
+
572
+ def assert_inconclusive(self, events: list[dict], reason: str | None = None) -> None:
573
+ failure = PARSER.audit_codex(self.args, events)
574
+ self.assertIsNotNone(failure)
575
+ self.assertEqual(failure["status"], "inconclusive")
576
+ if reason is not None:
577
+ self.assertEqual(failure["reason_code"], reason)
578
+
579
+ def test_codex_packet_accepts_verified_read(self) -> None:
580
+ failure = PARSER.audit_codex(
581
+ self.args, self.completed_events(self.read_event())
582
+ )
583
+
584
+ self.assertIsNone(failure)
585
+
586
+ def test_codex_packet_accepts_verified_search(self) -> None:
587
+ event = self.read_event()
588
+ arguments = {"query": "candidate", "byte_offset": 0, "limit": 2}
589
+ response = kimi_packet_mcp.search_packet(self.packet, self.digest, arguments)
590
+ self.assertFalse(response["isError"])
591
+ event["item"].update(tool="search_packet", arguments=arguments)
592
+ event["item"]["result"]["content"] = response["content"]
593
+
594
+ failure = PARSER.audit_codex(self.args, self.completed_events(event))
595
+
596
+ self.assertIsNone(failure)
597
+
598
+ def test_codex_packet_accepts_started_updated_completed_read(self) -> None:
599
+ completed = self.read_event()
600
+ started = copy.deepcopy(completed)
601
+ started["type"] = "item.started"
602
+ started["item"].update(status="in_progress", result=None)
603
+ updated = copy.deepcopy(started)
604
+ updated["type"] = "item.updated"
605
+
606
+ failure = PARSER.audit_codex(
607
+ self.args, self.completed_events(started, updated, completed)
608
+ )
609
+
610
+ self.assertIsNone(failure)
611
+
612
+ def test_codex_inline_accepts_completion_without_tools(self) -> None:
613
+ self.args.packet = None
614
+ self.args.packet_sha256 = None
615
+
616
+ self.assertIsNone(PARSER.audit_codex(self.args, self.completed_events()))
617
+
618
+ def test_codex_inline_rejects_packet_tool_without_binding(self) -> None:
619
+ self.args.packet = None
620
+ self.args.packet_sha256 = None
621
+
622
+ self.assert_inconclusive(
623
+ self.completed_events(self.read_event()), "tool_boundary_violation"
624
+ )
625
+
626
+ def test_codex_packet_rejects_unapproved_tool_surfaces(self) -> None:
627
+ read = self.read_event()["item"]
628
+ items = (
629
+ {**read, "server": "other_packet_server"},
630
+ {**read, "tool": "write_packet"},
631
+ {**read, "tool": "read_file"},
632
+ {
633
+ "id": "shell-1",
634
+ "type": "command_execution",
635
+ "command": "printf synthetic",
636
+ "aggregated_output": "synthetic",
637
+ "status": "completed",
638
+ "exit_code": 0,
639
+ },
640
+ )
641
+ for item in items:
642
+ with self.subTest(item_type=item["type"], tool=item.get("tool")):
643
+ event = {"type": "item.completed", "item": item}
644
+ self.assert_inconclusive(
645
+ self.completed_events(event), "tool_boundary_violation"
646
+ )
647
+
648
+ def test_codex_packet_rejects_forged_read_and_search_results(self) -> None:
649
+ for tool, arguments in (
650
+ ("read_packet", {"byte_offset": 0, "max_bytes": 46_000}),
651
+ ("search_packet", {"query": "candidate", "byte_offset": 0, "limit": 2}),
652
+ ):
653
+ with self.subTest(tool=tool):
654
+ event = self.read_event()
655
+ event["item"].update(tool=tool, arguments=arguments)
656
+ event["item"]["result"]["content"] = [
657
+ {"type": "text", "text": "forged packet contents"}
658
+ ]
659
+ self.assert_inconclusive(self.completed_events(event))
660
+
661
+ def test_codex_packet_rejects_changed_digest_bound_file(self) -> None:
662
+ event = self.read_event()
663
+ self.packet.write_text("Candidate\n", encoding="utf-8")
664
+
665
+ self.assert_inconclusive(self.completed_events(event))
666
+
667
+ def test_codex_packet_rejects_missing_read_completion(self) -> None:
668
+ event = self.read_event()
669
+ event["type"] = "item.started"
670
+ event["item"].update(status="in_progress", result=None)
671
+
672
+ self.assert_inconclusive(self.completed_events(event))
673
+
674
+ def test_codex_packet_rejects_malformed_and_extra_arguments(self) -> None:
675
+ cases = (
676
+ ("read_packet", None),
677
+ ("read_packet", []),
678
+ ("read_packet", {"byte_offset": 0}),
679
+ ("read_packet", {"byte_offset": 0, "max_bytes": 46_000, "path": "other"}),
680
+ ("search_packet", {"query": "candidate", "limit": 2}),
681
+ (
682
+ "search_packet",
683
+ {"query": "candidate", "byte_offset": 0, "limit": 2, "path": "other"},
684
+ ),
685
+ )
686
+ for tool, arguments in cases:
687
+ with self.subTest(tool=tool, arguments=arguments):
688
+ event = self.read_event()
689
+ event["item"].update(tool=tool, arguments=arguments)
690
+ self.assert_inconclusive(self.completed_events(event))
691
+
692
+ def test_codex_packet_accepts_verified_argument_error_then_retry(self) -> None:
693
+ retry = self.read_event()
694
+ retry["item"]["id"] = "read-retry"
695
+ error = self.read_event()
696
+ error["item"]["arguments"] = {"byte_offset": 11, "max_bytes": 46_000}
697
+ error["item"]["result"]["content"] = [
698
+ {"type": "text", "text": "packet chunk starts beyond end"}
699
+ ]
700
+
701
+ failure = PARSER.audit_codex(self.args, self.completed_events(error, retry))
702
+
703
+ self.assertIsNone(failure)
704
+
705
+ def test_codex_packet_transport_failure_stays_inconclusive_after_retry(self) -> None:
706
+ for status, result, error in (
707
+ ("failed", None, {"message": "transport closed"}),
708
+ ("completed", None, {"message": "transport closed"}),
709
+ ("completed", None, None),
710
+ ):
711
+ with self.subTest(status=status, error=error):
712
+ failed = self.read_event()
713
+ failed["item"].update(status=status, result=result, error=error)
714
+ retry = self.read_event()
715
+ retry["item"]["id"] = "read-retry"
716
+ self.assert_inconclusive(self.completed_events(failed, retry))
717
+
718
+
719
+ class CompletionFindingDispositionTest(unittest.TestCase):
720
+ """Exercise real receipt and completion validation without external models."""
721
+
722
+ def setUp(self) -> None:
723
+ temporary = tempfile.TemporaryDirectory()
724
+ self.addCleanup(temporary.cleanup)
725
+ self.root = Path(temporary.name)
726
+ self.packet = self.root / "candidate.patch"
727
+ self.packet.write_text(
728
+ "diff --git a/x b/x\n--- a/x\n+++ b/x\n@@ -1 +1 @@\n-a\n+b\n",
729
+ encoding="utf-8",
730
+ )
731
+ self.plan = self.root / "review-plan.json"
732
+ self.write_json(self.plan, {
733
+ "intent": "Preserve bounded completion after source refutation.",
734
+ "acceptance": ["Every finding occurrence is resolved without adding review authority."],
735
+ "self_review": [
736
+ {"concern": concern,
737
+ "conclusion": f"The synthetic completion fixture preserves {concern} boundaries.",
738
+ "evidence_refs": ["fixture"]}
739
+ for concern in ("correctness", "safety", "failure_paths", "tests_evidence", "compatibility")
740
+ ],
741
+ "evidence": [{"id": "fixture", "result": "Synthetic exact-candidate completion and history fixture."}],
742
+ })
743
+ self.dispositions = self.root / "dispositions.json"
744
+ self.record_receipts()
745
+
746
+ def record_receipts(self, final_status: str = "findings", *,
747
+ duplicate_findings: bool = False, distinct_finding: bool = False) -> None:
748
+ self.final_review_status = final_status
749
+ self.duplicate_findings = duplicate_findings
750
+ self.distinct_finding = distinct_finding
751
+ self.provider_calls: list[list[str]] = []
752
+ self.profiles: list[dict] = []
753
+ self.receipt_paths = [self.root / "review.json", self.root / "challenge.json"]
754
+ self.receipts = []
755
+ for index, mode in enumerate(("review", "challenge"), 1):
756
+ extra = ["--review-chain-id", "source-refutation-fixture",
757
+ "--autonomous-review-index", str(index)]
758
+ if mode == "challenge":
759
+ extra.extend(["--challenge-index", "1", "--focus", "source-boundary",
760
+ "--prior-review-result-file", str(self.receipt_paths[0])])
761
+ code, result = self.invoke(mode, extra)
762
+ self.assertEqual(code, 0, result)
763
+ self.assertEqual(result["status"], "findings" if mode == "review" else final_status, result)
764
+ self.receipts.append(result)
765
+ self.write_json(self.receipt_paths[index - 1], result)
766
+ self.original_receipt_bytes = [path.read_bytes() for path in self.receipt_paths]
767
+ self.refresh_dispositions()
768
+
769
+ @staticmethod
770
+ def write_json(path: Path, value: object) -> None:
771
+ path.write_text(json.dumps(value, ensure_ascii=False, sort_keys=True), encoding="utf-8")
772
+
773
+ @staticmethod
774
+ def digest(path: Path) -> str:
775
+ return hashlib.sha256(path.read_bytes()).hexdigest()
776
+
777
+ @staticmethod
778
+ def finding_digest(finding: dict) -> str:
779
+ canonical = json.dumps(finding, ensure_ascii=False, sort_keys=True, separators=(",", ":"))
780
+ return hashlib.sha256(canonical.encode("utf-8")).hexdigest()
781
+
782
+ def wrapper_result(self, command: list[str], **kwargs: object) -> subprocess.CompletedProcess:
783
+ # Only the provider boundary is replaced. Packet, profile, chain and
784
+ # completion validation execute normally; no subprocess is launched.
785
+ self.assertTrue(command[0].endswith("kimi_review.sh"), command[0])
786
+ self.provider_calls.append(list(command))
787
+ profile_path = Path(command[command.index("--review-profile-file") + 1])
788
+ profile = json.loads(profile_path.read_text(encoding="utf-8"))
789
+ self.profiles.append(profile)
790
+ mode = command[command.index("--mode") + 1]
791
+ payload = {
792
+ "reviewer": "kimi", "reviewer_family": "moonshot", "provider": "kimi-cli",
793
+ "model": "synthetic", "mode": mode, "status": "findings",
794
+ "native_skill_binding": "established",
795
+ "concern_results": [
796
+ {"concern": item["id"], "conclusion": f"Checked {item['id']} against the frozen candidate."}
797
+ for item in profile["required_concerns"]
798
+ ],
799
+ "findings": [{"severity": "P1", "file": "x", "line": 1,
800
+ "failure_path": "Synthetic missing guard before an action.",
801
+ "smallest_fix": "Check the guard before the synthetic action."}],
802
+ }
803
+ if mode == "challenge" and self.final_review_status == "passed":
804
+ payload.update(status="passed", findings=[])
805
+ elif self.duplicate_findings:
806
+ payload["findings"].append(dict(reversed(list(payload["findings"][0].items()))))
807
+ if self.distinct_finding:
808
+ payload["findings"].append({**payload["findings"][0],
809
+ "failure_path": "A distinct synthetic failure on the same line."})
810
+ return subprocess.CompletedProcess(command, 0, json.dumps(payload).encode("utf-8"), b"")
811
+
812
+ def arguments(self, mode: str) -> list[str]:
813
+ return ["--mode", mode, "--cwd", str(self.root), "--diff-file", str(self.packet),
814
+ "--implementer-family", "openai", "--review-plan-file", str(self.plan),
815
+ "--challenge-budget", "1"]
816
+
817
+ def invoke(self, mode: str, extra: list[str]) -> tuple[int, dict]:
818
+ output = io.StringIO()
819
+ with (mock.patch.object(review_gate, "run", side_effect=self.wrapper_result),
820
+ mock.patch.object(review_gate, "client_order", return_value=["kimi"]),
821
+ contextlib.redirect_stdout(output)):
822
+ code = review_gate.main(self.arguments(mode) + extra)
823
+ return code, json.loads(output.getvalue())
824
+
825
+ def refresh_dispositions(self) -> dict:
826
+ hashes = [self.digest(path) for path in self.receipt_paths]
827
+ receipts = [json.loads(path.read_text(encoding="utf-8")) for path in self.receipt_paths]
828
+ manifest = {
829
+ "schema_version": 1,
830
+ "candidate_sha256": receipts[-1]["candidate_sha256"],
831
+ "review_result_sha256": hashes,
832
+ "dispositions": [
833
+ {"receipt_sha256": receipt_hash,
834
+ "finding_sha256": self.finding_digest(finding),
835
+ "disposition": "source_refuted",
836
+ "evidence": ["Synthetic source x:1 checks the guard before the action; the reported path is unreachable."]}
837
+ for receipt_hash, receipt in zip(hashes, receipts)
838
+ for finding in {self.finding_digest(item): item for item in receipt["findings"]}.values()
839
+ ],
840
+ }
841
+ self.write_json(self.dispositions, manifest)
842
+ return manifest
843
+
844
+ def complete(self, *, include_prior: bool = True) -> tuple[int, dict]:
845
+ extra = ["--completion-review-result-file", str(self.receipt_paths[-1]),
846
+ "--finding-dispositions-file", str(self.dispositions)]
847
+ if include_prior:
848
+ extra.extend(["--prior-review-result-file", str(self.receipt_paths[0])])
849
+ result = self.invoke("complete", extra)
850
+ self.assertEqual(len(self.provider_calls), 2, "completion must never call a reviewer")
851
+ return result
852
+
853
+ def assert_rejected(self, *, include_prior: bool = True) -> dict:
854
+ code, result = self.complete(include_prior=include_prior)
855
+ self.assertEqual(code, 2, result)
856
+ self.assertEqual(result["status"], "inconclusive", result)
857
+ self.assertTrue(result["completion_gated"], result)
858
+ return result
859
+
860
+ def test_source_refuted_findings_validate_bound_dispositions(self) -> None:
861
+ # Keep the narrower validator exercised independently of CLI parsing.
862
+ args = review_gate.build_parser().parse_args(self.arguments("complete") + [
863
+ "--completion-review-result-file", str(self.receipt_paths[-1])])
864
+ args.finding_dispositions_file = str(self.dispositions)
865
+ args.prior_review_result_file = [str(self.receipt_paths[0])]
866
+ result_hash, prior, metadata = review_gate.validate_finding_dispositions(
867
+ args, self.receipts[-1]["candidate_sha256"], self.profiles[-1])
868
+ self.assertEqual(result_hash, self.digest(self.receipt_paths[-1]))
869
+ self.assertEqual(prior["findings"], self.receipts[-1]["findings"])
870
+ self.assertEqual(metadata["completion_basis"], "source_refuted_findings")
871
+
872
+ def test_two_refuted_rounds_complete_without_new_review_or_history_reset(self) -> None:
873
+ code, result = self.complete()
874
+ self.assertEqual(code, 0, result)
875
+ self.assertEqual(result["status"], "passed")
876
+ self.assertFalse(result["completion_gated"])
877
+ self.assertEqual(result["completion_basis"], "source_refuted_findings")
878
+ self.assertEqual(result["finding_dispositions_sha256"], self.digest(self.dispositions))
879
+ manifest = json.loads(self.dispositions.read_text(encoding="utf-8"))
880
+ # A repeated finding in a later round is a separate occurrence even
881
+ # when its canonical finding hash is identical.
882
+ self.assertEqual(len({item["finding_sha256"] for item in manifest["dispositions"]}), 1)
883
+ self.assertEqual(len({item["receipt_sha256"] for item in manifest["dispositions"]}), 2)
884
+ self.assertEqual(result["resolved_finding_occurrences"], [
885
+ {key: item[key] for key in ("receipt_sha256", "finding_sha256")}
886
+ for item in manifest["dispositions"]])
887
+ for field in ("review_chain_id", "review_scope_sha256", "prior_review_result_sha256",
888
+ "prior_challenge_focuses", "autonomous_review_budget", "autonomous_review_index",
889
+ "autonomous_reviews_remaining"):
890
+ self.assertEqual(result[field], self.receipts[-1][field], field)
891
+ self.assertEqual(result["autonomous_review_budget"], 2)
892
+ self.assertEqual(result["autonomous_review_index"], 2)
893
+ self.assertEqual(result["autonomous_reviews_remaining"], 0)
894
+ self.assertFalse(result["autonomous_review_allowed"])
895
+ self.assertEqual(result["next_action"], "complete")
896
+ self.assertEqual([path.read_bytes() for path in self.receipt_paths], self.original_receipt_bytes)
897
+
898
+ def test_dispositions_file_is_required_and_must_exist(self) -> None:
899
+ code, result = self.invoke("complete", [
900
+ "--completion-review-result-file", str(self.receipt_paths[-1])])
901
+ self.assertEqual(code, 2, result)
902
+ self.assertTrue(result["completion_gated"], result)
903
+ self.dispositions.unlink()
904
+ self.assert_rejected()
905
+
906
+ def test_identical_findings_share_identity_without_rewriting_receipts(self) -> None:
907
+ for final_status in ("findings", "passed"):
908
+ with self.subTest(final_status=final_status):
909
+ self.record_receipts(final_status, duplicate_findings=True)
910
+ self.assertEqual(len(self.receipts[0]["findings"]), 2)
911
+ code, result = self.complete()
912
+ self.assertEqual(code, 0, result)
913
+ self.assertEqual(len(result["resolved_finding_occurrences"]),
914
+ 2 if final_status == "findings" else 1)
915
+ self.assertEqual([path.read_bytes() for path in self.receipt_paths],
916
+ self.original_receipt_bytes)
917
+
918
+ def test_duplicate_identity_never_hides_a_distinct_finding_or_round(self) -> None:
919
+ self.record_receipts(duplicate_findings=True, distinct_finding=True)
920
+ manifest = self.refresh_dispositions()
921
+ self.assertEqual([len(row["findings"]) for row in self.receipts], [3, 3])
922
+ self.assertEqual(len(manifest["dispositions"]), 4)
923
+ code, result = self.complete()
924
+ self.assertEqual(code, 0, result)
925
+ self.assertEqual(len(result["resolved_finding_occurrences"]), 4)
926
+ for missing in range(4):
927
+ with self.subTest(missing_occurrence=missing):
928
+ incomplete = copy.deepcopy(manifest)
929
+ del incomplete["dispositions"][missing]
930
+ self.write_json(self.dispositions, incomplete)
931
+ self.assert_rejected()
932
+
933
+ def test_passed_final_still_requires_valid_historical_finding_dispositions(self) -> None:
934
+ self.record_receipts(final_status="passed")
935
+ code, result = self.complete()
936
+ self.assertEqual(code, 0, result)
937
+ self.assertEqual(result["completion_basis"], "source_refuted_findings")
938
+ self.assertEqual(len(result["resolved_finding_occurrences"]), 1)
939
+ self.assertEqual(result["resolved_finding_occurrences"][0]["receipt_sha256"],
940
+ self.digest(self.receipt_paths[0]))
941
+ manifest = self.refresh_dispositions()
942
+ manifest["dispositions"][0]["disposition"] = "accepted_risk"
943
+ self.write_json(self.dispositions, manifest)
944
+ rejected = self.assert_rejected()
945
+ self.assertEqual(rejected["next_action"], "resolve_review_findings")
946
+ manifest["dispositions"] = []
947
+ self.write_json(self.dispositions, manifest)
948
+ self.assert_rejected()
949
+
950
+ def test_every_historical_finding_occurrence_must_be_covered_once(self) -> None:
951
+ original = self.refresh_dispositions()
952
+ for missing in (0, 1):
953
+ with self.subTest(missing_round=missing):
954
+ manifest = copy.deepcopy(original)
955
+ del manifest["dispositions"][missing]
956
+ self.write_json(self.dispositions, manifest)
957
+ self.assert_rejected()
958
+ manifest = copy.deepcopy(original)
959
+ manifest["dispositions"].append(copy.deepcopy(manifest["dispositions"][0]))
960
+ self.write_json(self.dispositions, manifest)
961
+ self.assert_rejected()
962
+
963
+ def test_disposition_candidate_receipt_and_finding_hashes_are_bound(self) -> None:
964
+ original = self.refresh_dispositions()
965
+ for field in ("candidate_sha256", "review_result_sha256", "receipt_sha256", "finding_sha256"):
966
+ with self.subTest(field=field):
967
+ manifest = copy.deepcopy(original)
968
+ if field == "review_result_sha256":
969
+ manifest[field][0] = "f" * 64
970
+ elif field == "candidate_sha256":
971
+ manifest[field] = "f" * 64
972
+ else:
973
+ manifest["dispositions"][0][field] = "f" * 64
974
+ self.write_json(self.dispositions, manifest)
975
+ self.assert_rejected()
976
+
977
+ def test_unresolved_or_accepted_risk_cannot_complete(self) -> None:
978
+ original = self.refresh_dispositions()
979
+ for disposition in ("unresolved", "accepted_risk", "accepted_tradeoff", "needs_human_decision"):
980
+ with self.subTest(disposition=disposition):
981
+ manifest = copy.deepcopy(original)
982
+ manifest["dispositions"][0]["disposition"] = disposition
983
+ self.write_json(self.dispositions, manifest)
984
+ self.assert_rejected()
985
+
986
+ def test_source_refutation_requires_bounded_nonempty_evidence(self) -> None:
987
+ original = self.refresh_dispositions()
988
+ for evidence in (None, [], [""], [" \n"], "source x:1", [123], ["x" * 100_000]):
989
+ with self.subTest(evidence_type=type(evidence).__name__, size=len(str(evidence))):
990
+ manifest = copy.deepcopy(original)
991
+ if evidence is None:
992
+ del manifest["dispositions"][0]["evidence"]
993
+ else:
994
+ manifest["dispositions"][0]["evidence"] = evidence
995
+ self.write_json(self.dispositions, manifest)
996
+ self.assert_rejected()
997
+
998
+ def test_evidence_strings_match_downstream_normalization_and_text_boundary(self) -> None:
999
+ original = self.refresh_dispositions()
1000
+ for text in (" leading space", "trailing space ", "line\nbreak", "tab\tinside",
1001
+ "NUL\0inside", "DEL\x7finside", "C1\x85inside", "line\u2028separator",
1002
+ "paragraph\u2029separator", "zero\u200bwidth", "bidi\u202econtrol", "bad\ud800unicode"):
1003
+ with self.subTest(text=ascii(text)):
1004
+ manifest = copy.deepcopy(original)
1005
+ manifest["dispositions"][0]["evidence"] = [text]
1006
+ # Escaped JSON can encode an unpaired surrogate even though the
1007
+ # decoded string cannot be encoded as valid UTF-8 text.
1008
+ self.dispositions.write_text(json.dumps(manifest), encoding="utf-8")
1009
+ self.assert_rejected()
1010
+
1011
+ def test_changed_candidate_cannot_use_previous_refutations(self) -> None:
1012
+ self.packet.write_text(self.packet.read_text(encoding="utf-8").replace("+b\n", "+c\n"), encoding="utf-8")
1013
+ self.assert_rejected()
1014
+
1015
+ def test_stale_historical_candidate_cannot_use_current_refutations(self) -> None:
1016
+ historical = copy.deepcopy(self.receipts[0])
1017
+ historical["candidate_sha256"] = historical["packet_sha256"] = "f" * 64
1018
+ self.write_json(self.receipt_paths[0], historical)
1019
+ final = copy.deepcopy(self.receipts[-1])
1020
+ final["prior_review_result_sha256"][0] = self.digest(self.receipt_paths[0])
1021
+ self.write_json(self.receipt_paths[-1], final)
1022
+ # Refresh every outer receipt reference so stale candidate rejection
1023
+ # cannot be attributed to an incidental raw-receipt hash mismatch.
1024
+ self.refresh_dispositions()
1025
+ self.assert_rejected()
1026
+
1027
+ def test_inconclusive_receipt_cannot_be_refuted_into_a_pass(self) -> None:
1028
+ for index in (0, 1):
1029
+ with self.subTest(round=index):
1030
+ for path, raw in zip(self.receipt_paths, self.original_receipt_bytes):
1031
+ path.write_bytes(raw)
1032
+ receipt = json.loads(self.receipt_paths[index].read_text(encoding="utf-8"))
1033
+ receipt["status"] = "inconclusive"
1034
+ self.write_json(self.receipt_paths[index], receipt)
1035
+ if index == 0:
1036
+ final = copy.deepcopy(self.receipts[-1])
1037
+ final["prior_review_result_sha256"][0] = self.digest(self.receipt_paths[0])
1038
+ self.write_json(self.receipt_paths[-1], final)
1039
+ self.refresh_dispositions()
1040
+ self.assert_rejected()
1041
+
1042
+ def test_forged_or_missing_history_cannot_complete(self) -> None:
1043
+ self.assert_rejected(include_prior=False)
1044
+ receipt = copy.deepcopy(self.receipts[-1])
1045
+ receipt["prior_review_result_sha256"][0] = "f" * 64
1046
+ self.write_json(self.receipt_paths[-1], receipt)
1047
+ self.refresh_dispositions()
1048
+ self.assert_rejected()
1049
+
1050
+ def test_original_receipt_metadata_must_match_the_external_round(self) -> None:
1051
+ for index in (0, 1):
1052
+ for field in ("review_state", "human_decision_required", "gate_required", "gate_triggers"):
1053
+ with self.subTest(round=index, field=field):
1054
+ for path, raw in zip(self.receipt_paths, self.original_receipt_bytes):
1055
+ path.write_bytes(raw)
1056
+ receipt = copy.deepcopy(self.receipts[index])
1057
+ if field == "review_state":
1058
+ receipt[field] = "reviewed"
1059
+ elif field == "human_decision_required":
1060
+ receipt[field] = not receipt[field]
1061
+ elif field == "gate_required":
1062
+ receipt["self_review_gate"]["required"] = False
1063
+ else:
1064
+ receipt["self_review_gate"]["required_triggers"] = []
1065
+ self.write_json(self.receipt_paths[index], receipt)
1066
+ if index == 0:
1067
+ final = copy.deepcopy(self.receipts[-1])
1068
+ final["prior_review_result_sha256"][0] = self.digest(self.receipt_paths[0])
1069
+ self.write_json(self.receipt_paths[-1], final)
1070
+ self.refresh_dispositions()
1071
+ self.assert_rejected()
1072
+
1073
+ def test_receipt_order_is_part_of_the_disposition_binding(self) -> None:
1074
+ manifest = self.refresh_dispositions()
1075
+ manifest["review_result_sha256"].reverse()
1076
+ self.write_json(self.dispositions, manifest)
1077
+ self.assert_rejected()
1078
+
1079
+
508
1080
  if __name__ == "__main__":
509
1081
  unittest.main()
@@ -5,7 +5,9 @@ description: bug / 报错 / test 挂了 / 线上问题 / 接口变慢·性能退
5
5
 
6
6
  # Defect Diagnosis
7
7
 
8
- Use this skill for the full defect discipline: diagnose the immediate failure, fix it with evidence, then decide whether root-cause prevention should update a product, architecture, development, testing, release, or tooling skill.
8
+ Diagnose and fix from evidence; route prevention to product, architecture, development, testing, release or tooling.
9
+
10
+ - Code/test changes require self-checks; invoke `code-review` automatically before completion.
9
11
 
10
12
  ## Non-Negotiable Rules
11
13
 
@@ -5,7 +5,7 @@ description: Use when implementing, modifying, scaffolding, generating, or testi
5
5
 
6
6
  # Go Microservice Dev
7
7
 
8
- Use this for implementation of new backend products and services. It should adapt to the repo in front of you, but the workflow is independent of any prior codebase.
8
+ Use this for implementation of new backend products and services. It should adapt to the repo in front of you, but the workflow is independent of any prior codebase. After code/test edits, self-check and invoke `code-review` automatically before completion.
9
9
 
10
10
  ## Skill Routing
11
11
 
@@ -7,6 +7,8 @@ description: Use when designing, implementing, reviewing, debugging, or operatin
7
7
 
8
8
  Use this for product backend work that calls, hosts, evaluates, or operates LLM and inference systems. Keep the skill generic: extract reusable mechanics only, not business-specific prompts, datasets, provider names, repository paths, or domain nouns.
9
9
 
10
+ - Code/test changes require self-checks; invoke `code-review` automatically before completion.
11
+
10
12
  ## Skill Routing
11
13
 
12
14
  - Use this skill for LLM gateway/client design, model registry, prompt versioning, agent/tool orchestration, streaming APIs, fallback, token/cost accounting, evals, replay, shadow comparison, batch inference, and inference observability.
@@ -5,7 +5,7 @@ description: "小程序 / Taro / 微信小程序 / 支付宝小程序 / 抖音
5
5
 
6
6
  # Miniapp Product Dev
7
7
 
8
- Use this skill for mini-program client engineering and platform delivery. It covers product-facing miniapp work across WeChat, Alipay, Douyin/TikTok, Baidu, and similar host platforms. It does not own general product strategy, backend service architecture, or visual design rules.
8
+ Mini-program client engineering and product-facing delivery across WeChat, Alipay, Douyin/TikTok, Baidu and similar hosts; excludes product strategy, backend architecture and visual design. After code/test edits, self-check and invoke `code-review` automatically before completion.
9
9
 
10
10
  ## Framework Scope
11
11
 
@@ -57,7 +57,7 @@ This skill is about **using AI agents / subagents to execute work** — delegati
57
57
  5. **Value**: the task is worth that premium.
58
58
  If independence, value, or breadth is unclear, start with one focused agent or sequential delegation; use parallel multi-agent dispatch when those checks are explicitly satisfied.
59
59
  - Model tier per dispatch is an explicit decision, not an inherited accident. On hosts that inherit the session model for unnamed dispatches (a common default — Claude-family harnesses behave this way; verify yours), an unnamed model is often the most capable and most expensive tier, so a high-volume fan-out silently puts every worker and reviewer on the top tier. The observable triggers are a fan-out — multiple dispatches (workers/reviewers) in one delivery — and any high-risk dispatch: there, name the tier per dispatch and choose by judgment complexity and risk, not token price alone — the cheapest tier routinely takes 2–3× the turns on multi-step work and costs more overall, so use a mid-tier floor for reviewers and for implementers working from prose descriptions; reserve the cheapest tier for transcription-plus-tests tasks (the plan text already contains the code to write) and single-file mechanical fixes; put architecture/design judgment, security/authority/tenant-isolation/data-loss review, and the final whole-scope review on the most capable tier — review tier scales with the diff's size, complexity, and risk, and a high-risk review never silently inherits a cheap session default even as a single dispatch (tier principle: `../skill-extraction-workflow/references/harness-patterns-and-eval.md`). Record the decision as a brief field alongside `required_skills`: `model_tier: <tier>` or `model_tier: host-default (<reason: single low-risk dispatch | no host model selection>)` — an absent field is an unmade decision, not a default, and the field is bookkeeping only until the dispatch call actually passes the matching model selector (verify the effective model where the host exposes it; a field/selector mismatch is a defect, not a recorded decision).
60
- - Stop and escalate when the plan is unclear, a dependency is missing, verification fails repeatedly (same error ~3 times — identical retries, not new findings from successive review rounds), or an agent returns unsupported claims. An escalation message must carry five fields, or it is just "stuck": the specific blocker, the attempts made and their results, the current state (diff / commits / workspace), the safest next action for the human to align on, and whether a lower-risk part can continue meanwhile.
60
+ - Stop and escalate when unclear direction, unavailable dependencies, unknown completion state or missing authority remains after bounded remediation. Repeated identical verification failures (~3 times) require a status report and method/evidence checkpoint, not renewed task permission; stop identical retries and continue necessary work within existing scope, respecting explicit user limits. Escalation must name the blocker, attempts and results, current diff/commits/workspace, safest next action, and lower-risk work that can continue.
61
61
 
62
62
  ## Execution Flow
63
63