codex-orchestrator 2.0.1 → 2.0.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (211) hide show
  1. package/CHANGELOG.md +22 -0
  2. package/README.md +12 -9
  3. package/dist/src/index.d.ts +10 -0
  4. package/dist/src/index.d.ts.map +1 -1
  5. package/dist/src/index.js +5 -0
  6. package/dist/src/index.js.map +1 -1
  7. package/dist/src/v2/acceptance-proof.d.ts +3 -0
  8. package/dist/src/v2/acceptance-proof.d.ts.map +1 -1
  9. package/dist/src/v2/acceptance-proof.js +2 -8
  10. package/dist/src/v2/acceptance-proof.js.map +1 -1
  11. package/dist/src/v2/adapters/gh-issue-adapter.d.ts +5 -3
  12. package/dist/src/v2/adapters/gh-issue-adapter.d.ts.map +1 -1
  13. package/dist/src/v2/adapters/gh-issue-adapter.js +63 -7
  14. package/dist/src/v2/adapters/gh-issue-adapter.js.map +1 -1
  15. package/dist/src/v2/adapters/issues.d.ts +16 -2
  16. package/dist/src/v2/adapters/issues.d.ts.map +1 -1
  17. package/dist/src/v2/adapters/issues.js +15 -5
  18. package/dist/src/v2/adapters/issues.js.map +1 -1
  19. package/dist/src/v2/adapters/mission-coordinator-lock.d.ts +1 -0
  20. package/dist/src/v2/adapters/mission-coordinator-lock.d.ts.map +1 -1
  21. package/dist/src/v2/adapters/mission-coordinator-lock.js +5 -1
  22. package/dist/src/v2/adapters/mission-coordinator-lock.js.map +1 -1
  23. package/dist/src/v2/candidate-cli.d.ts +4 -0
  24. package/dist/src/v2/candidate-cli.d.ts.map +1 -1
  25. package/dist/src/v2/candidate-cli.js +26 -11
  26. package/dist/src/v2/candidate-cli.js.map +1 -1
  27. package/dist/src/v2/cli-contract.d.ts +1 -1
  28. package/dist/src/v2/cli-contract.d.ts.map +1 -1
  29. package/dist/src/v2/cli-contract.js +10 -0
  30. package/dist/src/v2/cli-contract.js.map +1 -1
  31. package/dist/src/v2/code-review-report.d.ts +66 -0
  32. package/dist/src/v2/code-review-report.d.ts.map +1 -0
  33. package/dist/src/v2/code-review-report.js +259 -0
  34. package/dist/src/v2/code-review-report.js.map +1 -0
  35. package/dist/src/v2/codex-process.d.ts +8 -1
  36. package/dist/src/v2/codex-process.d.ts.map +1 -1
  37. package/dist/src/v2/codex-process.js +11 -0
  38. package/dist/src/v2/codex-process.js.map +1 -1
  39. package/dist/src/v2/config.d.ts +2 -1
  40. package/dist/src/v2/config.d.ts.map +1 -1
  41. package/dist/src/v2/config.js +8 -3
  42. package/dist/src/v2/config.js.map +1 -1
  43. package/dist/src/v2/contained-report-operation.d.ts +100 -0
  44. package/dist/src/v2/contained-report-operation.d.ts.map +1 -0
  45. package/dist/src/v2/contained-report-operation.js +200 -0
  46. package/dist/src/v2/contained-report-operation.js.map +1 -0
  47. package/dist/src/v2/containment.d.ts +6 -0
  48. package/dist/src/v2/containment.d.ts.map +1 -1
  49. package/dist/src/v2/containment.js +40 -1
  50. package/dist/src/v2/containment.js.map +1 -1
  51. package/dist/src/v2/direct-delivery.d.ts +101 -0
  52. package/dist/src/v2/direct-delivery.d.ts.map +1 -0
  53. package/dist/src/v2/direct-delivery.js +547 -0
  54. package/dist/src/v2/direct-delivery.js.map +1 -0
  55. package/dist/src/v2/immutable-workflow-publisher.d.ts +40 -0
  56. package/dist/src/v2/immutable-workflow-publisher.d.ts.map +1 -0
  57. package/dist/src/v2/immutable-workflow-publisher.js +218 -0
  58. package/dist/src/v2/immutable-workflow-publisher.js.map +1 -0
  59. package/dist/src/v2/implementation-reviewer.d.ts +81 -0
  60. package/dist/src/v2/implementation-reviewer.d.ts.map +1 -0
  61. package/dist/src/v2/implementation-reviewer.js +157 -0
  62. package/dist/src/v2/implementation-reviewer.js.map +1 -0
  63. package/dist/src/v2/owner-control-lock.d.ts +41 -0
  64. package/dist/src/v2/owner-control-lock.d.ts.map +1 -0
  65. package/dist/src/v2/owner-control-lock.js +174 -0
  66. package/dist/src/v2/owner-control-lock.js.map +1 -0
  67. package/dist/src/v2/route-continuations.d.ts +32 -0
  68. package/dist/src/v2/route-continuations.d.ts.map +1 -0
  69. package/dist/src/v2/route-continuations.js +2 -0
  70. package/dist/src/v2/route-continuations.js.map +1 -0
  71. package/dist/src/v2/route-coordinator.d.ts +77 -0
  72. package/dist/src/v2/route-coordinator.d.ts.map +1 -0
  73. package/dist/src/v2/route-coordinator.js +370 -0
  74. package/dist/src/v2/route-coordinator.js.map +1 -0
  75. package/dist/src/v2/route-decision.d.ts +129 -0
  76. package/dist/src/v2/route-decision.d.ts.map +1 -0
  77. package/dist/src/v2/route-decision.js +400 -0
  78. package/dist/src/v2/route-decision.js.map +1 -0
  79. package/dist/src/v2/run-issue.d.ts +63 -2
  80. package/dist/src/v2/run-issue.d.ts.map +1 -1
  81. package/dist/src/v2/run-issue.js +906 -91
  82. package/dist/src/v2/run-issue.js.map +1 -1
  83. package/dist/src/v2/run-store.d.ts +25 -1
  84. package/dist/src/v2/run-store.d.ts.map +1 -1
  85. package/dist/src/v2/run-store.js +143 -3
  86. package/dist/src/v2/run-store.js.map +1 -1
  87. package/dist/src/v2/runtime-assets.d.ts +15 -13
  88. package/dist/src/v2/runtime-assets.d.ts.map +1 -1
  89. package/dist/src/v2/runtime-assets.js +263 -416
  90. package/dist/src/v2/runtime-assets.js.map +1 -1
  91. package/dist/src/v2/runtime.d.ts +14 -6
  92. package/dist/src/v2/runtime.d.ts.map +1 -1
  93. package/dist/src/v2/runtime.js +478 -56
  94. package/dist/src/v2/runtime.js.map +1 -1
  95. package/dist/src/v2/setup-cli.d.ts.map +1 -1
  96. package/dist/src/v2/setup-cli.js +1 -0
  97. package/dist/src/v2/setup-cli.js.map +1 -1
  98. package/dist/src/v2/setup-runtime.d.ts.map +1 -1
  99. package/dist/src/v2/setup-runtime.js +20 -72
  100. package/dist/src/v2/setup-runtime.js.map +1 -1
  101. package/dist/src/v2/setup.d.ts +4 -1
  102. package/dist/src/v2/setup.d.ts.map +1 -1
  103. package/dist/src/v2/setup.js +104 -1
  104. package/dist/src/v2/setup.js.map +1 -1
  105. package/dist/src/v2/spec-coordinator.d.ts +85 -0
  106. package/dist/src/v2/spec-coordinator.d.ts.map +1 -0
  107. package/dist/src/v2/spec-coordinator.js +88 -0
  108. package/dist/src/v2/spec-coordinator.js.map +1 -0
  109. package/dist/src/v2/spec-delivery.d.ts +143 -0
  110. package/dist/src/v2/spec-delivery.d.ts.map +1 -0
  111. package/dist/src/v2/spec-delivery.js +401 -0
  112. package/dist/src/v2/spec-delivery.js.map +1 -0
  113. package/dist/src/v2/triage-route.d.ts +68 -0
  114. package/dist/src/v2/triage-route.d.ts.map +1 -0
  115. package/dist/src/v2/triage-route.js +223 -0
  116. package/dist/src/v2/triage-route.js.map +1 -0
  117. package/dist/src/v2/waiting-human-coordinator.d.ts +49 -0
  118. package/dist/src/v2/waiting-human-coordinator.d.ts.map +1 -0
  119. package/dist/src/v2/waiting-human-coordinator.js +509 -0
  120. package/dist/src/v2/waiting-human-coordinator.js.map +1 -0
  121. package/dist/src/v2/waiting-human.d.ts +143 -0
  122. package/dist/src/v2/waiting-human.d.ts.map +1 -0
  123. package/dist/src/v2/waiting-human.js +408 -0
  124. package/dist/src/v2/waiting-human.js.map +1 -0
  125. package/dist/src/v2/workflow-assets.d.ts +90 -0
  126. package/dist/src/v2/workflow-assets.d.ts.map +1 -0
  127. package/dist/src/v2/workflow-assets.js +554 -0
  128. package/dist/src/v2/workflow-assets.js.map +1 -0
  129. package/docs/deep-dive.md +15 -8
  130. package/internal-workflow/docs/agents/artifact-review-loop.md +267 -0
  131. package/internal-workflow/docs/agents/bug-workflow-routing.md +24 -0
  132. package/internal-workflow/docs/agents/coding-skill-routing.md +203 -0
  133. package/internal-workflow/docs/agents/confidence-rubric.md +65 -0
  134. package/internal-workflow/docs/agents/contract-test-ledger.md +60 -0
  135. package/internal-workflow/docs/agents/implementation-review-loop.md +302 -0
  136. package/internal-workflow/docs/agents/review-gates.md +49 -0
  137. package/internal-workflow/docs/agents/review-protocol.md +170 -0
  138. package/internal-workflow/docs/agents/tool-usage.md +88 -0
  139. package/internal-workflow/manifest.json +1 -0
  140. package/internal-workflow/operations/acceptance-proof/SKILL.md +3 -0
  141. package/internal-workflow/operations/ambiguity-review/SKILL.md +3 -0
  142. package/internal-workflow/operations/cleanup-review/SKILL.md +3 -0
  143. package/internal-workflow/operations/code-review/SKILL.md +3 -0
  144. package/internal-workflow/operations/implementation/SKILL.md +3 -0
  145. package/internal-workflow/operations/spec-author/SKILL.md +3 -0
  146. package/internal-workflow/operations/spec-implementation/SKILL.md +3 -0
  147. package/internal-workflow/operations/spec-review/SKILL.md +3 -0
  148. package/internal-workflow/operations/triage/SKILL.md +3 -0
  149. package/internal-workflow/profiles/analyst_deep.toml +9 -0
  150. package/internal-workflow/profiles/implementer_deep.toml +9 -0
  151. package/internal-workflow/profiles/implementer_standard.toml +9 -0
  152. package/internal-workflow/profiles/proof_agent.toml +8 -0
  153. package/internal-workflow/profiles/researcher_standard.toml +9 -0
  154. package/internal-workflow/profiles/reviewer_deep.toml +9 -0
  155. package/internal-workflow/profiles/reviewer_fast.toml +9 -0
  156. package/internal-workflow/profiles/reviewer_standard.toml +9 -0
  157. package/internal-workflow/schemas/ambiguity-review-v1.json +1 -0
  158. package/internal-workflow/schemas/code-review-v1.json +1 -0
  159. package/internal-workflow/schemas/implementation-report-v1.json +1 -0
  160. package/internal-workflow/schemas/proof-report-v1.json +1 -0
  161. package/internal-workflow/schemas/spec-author-v1.json +1 -0
  162. package/internal-workflow/schemas/spec-review-v1.json +30 -0
  163. package/internal-workflow/schemas/triage-route-v1.json +1 -0
  164. package/internal-workflow/skills/acceptance-proof/agents/openai.yaml +6 -0
  165. package/internal-workflow/skills/agent-auto/agents/openai.yaml +6 -0
  166. package/internal-workflow/skills/cleanup-review/SKILL.md +84 -0
  167. package/internal-workflow/skills/cleanup-review/agents/openai.yaml +6 -0
  168. package/internal-workflow/skills/code-review/SKILL.md +257 -0
  169. package/internal-workflow/skills/code-review/agents/openai.yaml +4 -0
  170. package/internal-workflow/skills/code-review/references/bug-classes.md +56 -0
  171. package/internal-workflow/skills/code-review/references/framework-lenses.md +34 -0
  172. package/internal-workflow/skills/code-review/references/targeted-recipes.md +49 -0
  173. package/internal-workflow/skills/codebase-design/DEEPENING.md +35 -0
  174. package/internal-workflow/skills/codebase-design/DESIGN-IT-TWICE.md +50 -0
  175. package/internal-workflow/skills/codebase-design/SKILL.md +82 -0
  176. package/internal-workflow/skills/codebase-design/agents/openai.yaml +6 -0
  177. package/internal-workflow/skills/diagnosing-bugs/SKILL.md +138 -0
  178. package/internal-workflow/skills/diagnosing-bugs/agents/openai.yaml +6 -0
  179. package/internal-workflow/skills/diagnosing-bugs/scripts/hitl-loop.template.sh +41 -0
  180. package/internal-workflow/skills/implementation-spec-maker/SKILL.md +93 -0
  181. package/internal-workflow/skills/implementation-spec-maker/agents/openai.yaml +6 -0
  182. package/internal-workflow/skills/implementation-spec-maker/references/source-modes.md +31 -0
  183. package/internal-workflow/skills/implementation-spec-maker/references/spec-template.md +146 -0
  184. package/internal-workflow/skills/implementation-spec-review/SKILL.md +211 -0
  185. package/internal-workflow/skills/implementation-spec-review/agents/openai.yaml +6 -0
  186. package/internal-workflow/skills/research/SKILL.md +107 -0
  187. package/internal-workflow/skills/research/agents/openai.yaml +6 -0
  188. package/internal-workflow/skills/small-task-implementer/SKILL.md +97 -0
  189. package/internal-workflow/skills/small-task-implementer/agents/openai.yaml +6 -0
  190. package/internal-workflow/skills/spec-implementer/SKILL.md +197 -0
  191. package/internal-workflow/skills/spec-implementer/agents/openai.yaml +6 -0
  192. package/internal-workflow/skills/tdd/SKILL.md +59 -0
  193. package/internal-workflow/skills/tdd/agents/openai.yaml +6 -0
  194. package/internal-workflow/skills/tdd/interface-design.md +31 -0
  195. package/internal-workflow/skills/tdd/mocking.md +59 -0
  196. package/internal-workflow/skills/tdd/refactoring.md +10 -0
  197. package/internal-workflow/skills/tdd/tests.md +77 -0
  198. package/internal-workflow/skills/triage/AGENT-BRIEF.md +192 -0
  199. package/internal-workflow/skills/triage/OUT-OF-SCOPE.md +101 -0
  200. package/internal-workflow/skills/triage/SKILL.md +134 -0
  201. package/internal-workflow/skills/triage/agents/openai.yaml +6 -0
  202. package/internal-workflow/skills/ui-evidence-proof/SKILL.md +123 -0
  203. package/internal-workflow/skills/ui-evidence-proof/agents/openai.yaml +6 -0
  204. package/package.json +6 -3
  205. /package/{internal-skills → internal-workflow/skills}/acceptance-proof/SKILL.md +0 -0
  206. /package/{internal-skills → internal-workflow/skills}/acceptance-proof/references/android.md +0 -0
  207. /package/{internal-skills → internal-workflow/skills}/acceptance-proof/references/browser.md +0 -0
  208. /package/{internal-skills → internal-workflow/skills}/acceptance-proof/references/ios.md +0 -0
  209. /package/{internal-skills → internal-workflow/skills}/acceptance-proof/tools/android-lease.mjs +0 -0
  210. /package/{internal-skills → internal-workflow/skills}/acceptance-proof/tools/ios-lease.mjs +0 -0
  211. /package/{internal-skills → internal-workflow/skills}/agent-auto/SKILL.md +0 -0
@@ -0,0 +1,65 @@
1
+ # Confidence Rubric
2
+
3
+ Use this rubric for coding skills that report findings, diagnose root causes, review specs, or decide whether an auto-fix is safe.
4
+
5
+ Do not use fake numeric precision unless a skill has a concrete scoring reason. Prefer `high`, `medium`, and `low` with evidence.
6
+
7
+ ## High Confidence
8
+
9
+ High confidence means the finding or diagnosis has direct evidence.
10
+
11
+ Requires:
12
+
13
+ - a concrete trigger path, code path, failing signal, or tool output
14
+ - a clear explanation of why existing guards do not prevent the issue
15
+ - no unresolved assumption that changes the conclusion
16
+
17
+ Allowed actions:
18
+
19
+ - report as a finding
20
+ - block execution when the issue is safety-critical or spec-critical
21
+ - auto-fix only when the fix is narrow, low-risk, local-patterned, and verifiable
22
+
23
+ ## Medium Confidence
24
+
25
+ Medium confidence means the issue is likely but one explicit assumption remains.
26
+
27
+ Requires:
28
+
29
+ - strong local evidence
30
+ - exactly what assumption remains
31
+ - what evidence would promote or demote the finding
32
+
33
+ Allowed actions:
34
+
35
+ - report as a likely issue, risk, or execution concern
36
+ - ask a targeted question when the unresolved assumption changes the fix
37
+ - do not auto-fix unless new evidence raises confidence to high
38
+
39
+ ## Low Confidence
40
+
41
+ Low confidence means the concern is plausible but not proven.
42
+
43
+ Requires:
44
+
45
+ - a clear label as uncertainty
46
+ - the missing evidence or verification gap
47
+
48
+ Allowed actions:
49
+
50
+ - present as a question, risk, or verification gap
51
+ - do not report as a proven bug
52
+ - do not auto-fix
53
+
54
+ ## Auto-Fix Gate
55
+
56
+ Auto-fix is allowed only when all are true:
57
+
58
+ - confidence is high
59
+ - root cause is clear
60
+ - fix is narrow and low-risk
61
+ - fix matches local project patterns
62
+ - verification is available, or the edit is syntax-checkable and obviously safe
63
+ - the change does not require a product decision
64
+
65
+ If any condition is missing, report the issue with evidence and stop before editing.
@@ -0,0 +1,60 @@
1
+ # Contract Test Ledger
2
+
3
+ Use this shared ledger for behavior-changing work where a passing happy-path test could still miss a contract defect. The ledger turns review-class risks into testable obligations before implementation.
4
+
5
+ ## When Required
6
+
7
+ Create or update a contract test ledger when the task changes any of these:
8
+
9
+ - API, DTO, schema, serialization, persistence, or externally visible response shape
10
+ - ordering, lifecycle events, state transitions, retries, idempotency, timeout, cancellation, or background jobs
11
+ - cache keys, invalidation, state merge precedence, fallback behavior, profile/global/mobile overrides, defaults, or feature flags
12
+ - evidence, trace, snapshot, audit, summary, aggregation, score, winner, or generated artifacts
13
+ - shared behavior read by multiple callers, tenants, users, groups, children, or projections
14
+
15
+ For narrow UI copy, docs-only, formatting, tests-only, or isolated styling changes, the ledger is not required.
16
+
17
+ ## Required Shape
18
+
19
+ Keep the ledger compact. Use one row per invariant:
20
+
21
+ ```markdown
22
+ ## Contract Test Ledger
23
+
24
+ | Invariant | Risk It Prevents | First Test / Proof | Status |
25
+ | --- | --- | --- | --- |
26
+ | <observable rule> | <real failure mode> | <exact test name/command or manual proof> | planned / red / green / blocked |
27
+ ```
28
+
29
+ Rules:
30
+
31
+ - The invariant must be observable through the public interface or the same seam real callers use.
32
+ - The risk must name the concrete bug class, not a vague "edge case".
33
+ - The first test/proof must fail before the fix unless the ledger records why a RED signal is impossible.
34
+ - `blocked` requires the missing seam, fixture, service, or decision that prevents proof.
35
+ - Keep the ledger current as implementation proceeds; do not backfill it only at the end.
36
+
37
+ ## Invariant Prompts
38
+
39
+ Ask the relevant subset before the first RED test:
40
+
41
+ - **Ordering:** What must happen before/after terminal events, snapshots, persistence writes, notifications, or cleanup?
42
+ - **Precedence:** Which source wins among user input, profile, mobile, global, server, cache, AI, fallback, default, `null`, `false`, `0`, and empty objects?
43
+ - **Threading:** Does each new field survive construction, normalization, cloning, retry, persistence reload, serialization, and every visible consumer?
44
+ - **Runtime contract:** Do validation, internal types, persistence schema, API response, and consumers agree on names, units, nullability, enum values, and date/object formats?
45
+ - **Retry/idempotency:** What changes on retry, and which snapshots, counters, streams, timestamps, writes, or side effects must be rebuilt instead of reused?
46
+ - **Determinism:** When sort keys, timestamps, scores, priorities, or winners tie, what stable tie-breaker makes output repeatable?
47
+ - **Evidence:** Which trace, audit, snapshot, Fresh-Context, summary, or generated artifact proves the behavior actually happened?
48
+ - **Partial failure:** If a dependency times out, throws, returns stale data, or fails after a side effect, what durable state remains and who repairs it?
49
+ - **Scope/cardinality:** Is data global, per-tenant, per-group, per-child, per-step, or per-item, and can one top-level field collapse multiple meaningful results?
50
+
51
+ ## Review Feedback Loop
52
+
53
+ When code review finds a real contract defect, add or update one ledger row before fixing it:
54
+
55
+ - `Invariant`: the rule the implementation violated
56
+ - `Risk It Prevents`: the observed review finding
57
+ - `First Test / Proof`: the regression test or proof that would have caught it
58
+ - `Status`: `red` before the fix, then `green` after verification
59
+
60
+ If no correct public seam exists for the regression test, record that as `blocked` and name the architecture/testability gap. Do not replace a missing seam with an implementation-detail test unless the task explicitly approves that tradeoff.
@@ -0,0 +1,302 @@
1
+ # Implementation Review Loop
2
+
3
+ Use this policy for every approved implementation-spec execution. This
4
+ target-specific Module owns implementation authority, durable Review State,
5
+ checkpoint/final topology, validation reuse, gate ordering, audit epochs, and
6
+ implementation outcome mapping. It applies
7
+ [`review-protocol.md`](review-protocol.md) for context transfer, Full/Closure
8
+ mechanics, defect lifecycle, no-progress, and the common result envelope.
9
+ `spec-implementer`, `cleanup-review`, and `code-review` are callers or Adapters;
10
+ they must not reproduce either Module. Deterministic tickets-orchestrator work
11
+ using its issue as authority remains outside this Module and uses direct TDD
12
+ plus repo review gates.
13
+
14
+ This policy does not replace tests, architecture checks, smoke tests, or Git
15
+ checkpoints. Those proofs remain independent evidence.
16
+
17
+ ## Interface
18
+
19
+ Conceptually, the executor uses:
20
+
21
+ ```text
22
+ review_implementation(
23
+ authority_artifact_kind,
24
+ authority_artifact_path,
25
+ review_profile,
26
+ current_revision,
27
+ checkpoint,
28
+ review_focus,
29
+ defect_ledger
30
+ ) -> review_outcome
31
+ ```
32
+
33
+ The outcome records:
34
+
35
+ - `outcome`: `Approved | Blocked | Waived`
36
+ - `authority_artifact_kind`: `approved-spec`
37
+ - `authority_artifact_path`: the sole artifact that stores review state
38
+ - `review_profile`: `simple | medium | high`
39
+ - completed review passes and pending launches
40
+ - review mode, reviewer/session identity, target revision, and assigned lenses
41
+ - stable defect IDs and their current status
42
+ - logical skill activations and their open/closed state
43
+ - mandatory final coverage still required
44
+
45
+ ## Authority
46
+
47
+ An approved implementation spec stores the sole `## Implementation Review
48
+ State`. Architecture RFCs, product PRDs, tickets, `ready-for-approval`
49
+ artifacts, inferred approval, and ordinary direct-ticket waves are ineligible.
50
+ Persist `authority_artifact_kind` and `authority_artifact_path` and never create
51
+ a second ledger in an upstream artifact, caller, ticket, or Adapter.
52
+
53
+ Lifecycle and proof updates to the selected artifact do not change its approved
54
+ status or substantive design. A substantive authority change requires the
55
+ normal artifact revision/review path before implementation continues.
56
+
57
+ ## Profile And Review Shape
58
+
59
+ Prefer the selected authority artifact's `review_profile`. If it is absent, use the same
60
+ evidence-based classification and hard escalators as
61
+ [`artifact-review-loop.md`](artifact-review-loop.md). Actual implementation
62
+ evidence may raise the profile but must not lower it.
63
+
64
+ Review profile selects mandatory lenses and independence. Protocol pass-count
65
+ semantics apply; parallel reviewer results remain separate passes even when they
66
+ reduce wall-clock time.
67
+
68
+ Select the reviewer role from the profile: `simple` uses `reviewer_fast`,
69
+ `medium` uses `reviewer_standard`, and `high` uses `reviewer_deep`. Root always
70
+ launches reviewer children; it never performs an implementation review inline.
71
+ An Adapter runs inline only inside its already assigned reviewer child.
72
+
73
+ A reviewer that fails before returning a usable result is recorded as failed
74
+ and closed, not as completed coverage. Escalation preserves completed coverage
75
+ and the stable Defect Ledger. A user pause, context compaction, slice commit,
76
+ worker replacement, or new turn also preserves them; none restarts the review
77
+ topology automatically.
78
+
79
+ Stop early when mandatory coverage and protocol clear-state requirements hold.
80
+
81
+ ## Review Planning
82
+
83
+ Before the first implementation reviewer, root creates a short Review Plan:
84
+
85
+ - current profile and required independent lenses
86
+ - only stable intermediate checkpoints and their required lenses; move an unstable checkpoint to final coverage when later slices touch the same files, owners, or contracts
87
+ - any separate cleanup requirement, which must name a concrete evidenced reason that cannot fit the final spec/standards lens
88
+ - final code-review lenses and minimum independent coverage
89
+ - reviewer lineages that own affected-lens Closure
90
+
91
+ Create one durable activation record for each logical skill invocation. Record
92
+ `activation_id`, skill, owner, opened/closed state, and resume rule. Review Full
93
+ and lineage-preserving Closure passes stay inside that review skill's activation;
94
+ TDD repair cycles stay inside the active TDD activation. Cleanup, code review,
95
+ TDD, and debugger activations never share an ID, and a continuation resumes an
96
+ ID only for the same skill and authorized flow.
97
+
98
+ ## Durable Review State
99
+
100
+ Do not create durable review state during implementation preflight. Immediately
101
+ before the first actual reviewer launch, persist the short Review Plan and
102
+ pending launch in the selected authority artifact under
103
+ `## Implementation Review State`. From that point onward this is the execution
104
+ ledger for review state; do not keep the authoritative history only in chat
105
+ context or a subagent summary.
106
+
107
+ Record at least:
108
+
109
+ - profile, completed pass count, and required coverage still outstanding
110
+ - current checkpoint/gate, review timing baseline, gate-local consecutive
111
+ Closure-wave count, and latest Closure wave ID
112
+ - planned mandatory final reviews and their lenses
113
+ - authority artifact kind/path and logical skill activation records
114
+ - each lineage ID, origin Full session, active session generation, Closure count,
115
+ rotation reason, live/timeout state, and `conclude_requested_at`
116
+ - any convergence audit epoch: trigger, triggering pass/wave/revision,
117
+ completion, dispositions, selected sessions, resume reason, and pass/wave/time
118
+ baselines used for its next rearm
119
+ - pending reviewer launches with launch ID, mode, lineage/session identity,
120
+ activation ID, target revision, checkpoint/gate, Closure wave ID when
121
+ applicable, assigned lenses, and start timestamp
122
+ - every completed review's mode, lineage/session identity, target revision,
123
+ checkpoint/gate, Closure wave ID, assigned lenses, start/end timestamps, and
124
+ outcome
125
+ - the stable Defect Ledger with transition history, reopen count, fixed revision,
126
+ verifying review, and any explicit risk acceptance
127
+
128
+ Write a pending launch before starting its reviewer. After the launch returns,
129
+ replace the pending record with either its usable completed result or a failed,
130
+ closed session record; only a usable result increments `review_passes`. A context
131
+ compaction, new turn, resumed task, or different root agent must reconcile every
132
+ pending launch with its recorded session before starting a replacement.
133
+
134
+ Update this section after each launch, usable reviewer result, repair batch,
135
+ closure, waiver, acceptance, reopen, or terminal outcome. A resumed executor
136
+ reconstructs accounting and lifecycle history from this persisted state. If the
137
+ state is missing or internally inconsistent after reviews began, return
138
+ `Blocked` until it is reconciled from available thread/session evidence; never
139
+ assume zero completed passes or silently replace an in-flight reviewer.
140
+
141
+ Plan mandatory final coverage before launching an intermediate review. Launch a
142
+ checkpoint only when its target is settled and later slices will not invalidate
143
+ the reviewed files, owners, or contracts. Otherwise move its lenses to final
144
+ coverage. Do not replace a required final lens with another fresh checkpoint
145
+ reviewer or a repeat broad audit.
146
+
147
+ The default shapes are:
148
+
149
+ - `simple`: validation only when policy does not require review; otherwise one
150
+ final Full review and affected-lens Closure only after repairs.
151
+ - `medium`: an explicit intermediate checkpoint may provide one required lens;
152
+ the final integrator covers every remaining lens, includes bounded cleanup in
153
+ spec/standards, and verifies its defects without a separate cleanup pass.
154
+ - `high`: use parallel independent tracks only for disjoint mandatory lenses;
155
+ the spec/standards track includes bounded cleanup, and affected-lens Closure
156
+ follows only after consolidated repairs. There is no separate cleanup pass by
157
+ default.
158
+
159
+ `code-review` uses one final reviewer covering both correctness and
160
+ spec/standards for `simple` and `medium`. For `high`, it uses two disjoint final
161
+ tracks unless earlier independent coverage already covered both axes and one
162
+ fresh final integrator receives their compact handoffs.
163
+
164
+ Before launching a fresh final reviewer, reconcile coverage on the settled
165
+ revision. If the latest usable Full or Closure covered every mandatory final
166
+ lens and left no open defect, mark final review complete and stop. Count cleanup
167
+ Closure only for the lenses explicitly assigned in the Review Plan; a
168
+ `cleanup-only` pass does not satisfy correctness or spec/standards coverage.
169
+
170
+ ## Review Capsule
171
+
172
+ Use the protocol capsule with these implementation fields:
173
+
174
+ - the unanswered implementation question and any prior coverage it invalidates
175
+ - authority artifact kind/path, profile, current revision, checkpoint, and exact diff command
176
+ - changed paths and assigned `Review Focus` lenses
177
+ - source-of-truth docs and relevant Contract Test Ledger rows
178
+ - compact validation results and known verification gaps
179
+
180
+ For Closure, map changed paths and tests into the protocol Revision Map.
181
+
182
+ ## Review Modes
183
+
184
+ Use protocol Full and Closure without redefining them. Implementation Closure
185
+ maps `affected_targets` to paths, tests, runtime contracts, and Review Focus
186
+ lenses. An already planned Full reviewer may verify a repair when its assigned
187
+ lenses cover it.
188
+
189
+ ## Defect Lifecycle
190
+
191
+ Use the canonical protocol ledger and lifecycle without local aliases.
192
+
193
+ Implementation proof-only gaps may use `planned-final-verification` only when an
194
+ already scheduled code-review lens owns the proof; they remain open until
195
+ independently verified and never re-enter cleanup. Artifact proof-contract gaps
196
+ reopen artifact review. A pre-existing adjacent issue is non-blocking only as an
197
+ `improvement` with `follow-up-improvement`.
198
+
199
+ ## Validation Evidence Reuse
200
+
201
+ Persist command/config identity, failure signature, target revision and changed
202
+ path/contract impact basis, secret-safe environment fingerprint, transitive
203
+ ownership/contract impact, and result. Reuse a known unrelated suite failure
204
+ only when every field matches; unknown environment or transitive impact fails
205
+ closed. Focused tests and every required check for a repair always rerun.
206
+
207
+ ## Gate Ordering
208
+
209
+ 1. Implement the slice and pass its tests/exit gate.
210
+ 2. At an explicit intermediate checkpoint, run the required targeted
211
+ `code-review` directly under the Review Plan.
212
+ 3. Repair one consolidated finding batch and use protocol Closure for the
213
+ affected lineages. An already planned Full reviewer may verify the repair when
214
+ its assigned lenses cover it.
215
+ At one gate, collect the usable results from all already-launched reviewers
216
+ before repairing, unless an immediate blocker invalidates the remaining
217
+ work. Do not turn individual findings into serial repair, validation, and
218
+ Closure micro-cycles. Repair compatible findings once, rerun each affected
219
+ validation once on the resulting revision, then launch one affected-lens
220
+ Closure wave.
221
+ 4. Continue implementation only when checkpoint blockers are verified or the
222
+ Review Plan explicitly assigns their verification to an already planned reviewer
223
+ without violating the checkpoint's safety purpose.
224
+ 5. After all implementation slices and validations settle, run one final code
225
+ review wave. `simple` and `medium` use one reviewer; `high` launches two
226
+ disjoint reviewer tracks in parallel. The spec/standards lens owns bounded
227
+ cleanup.
228
+
229
+ Intermediate code-review checkpoints do not run cleanup-review. A separate
230
+ cleanup pass is exceptional: run it only when the user, approved source, or repo
231
+ policy names a concrete evidenced simplification risk that cannot fit the final
232
+ spec/standards lens. `large` or `high` alone is not a reason. If an approved spec
233
+ names an intermediate cleanup checkpoint, return `Blocked` for spec revision.
234
+
235
+ Cleanup review runs at most once as a Full review for the whole spec. After its
236
+ findings are repaired, either use protocol Closure or give the final code
237
+ reviewer those stable defect IDs for verification. Never launch another Full
238
+ cleanup review over the repaired whole diff.
239
+
240
+ ## Convergence And Stop Rules
241
+
242
+ Apply protocol repair, no-progress, stop, and waiver semantics. The following
243
+ audit is implementation-specific.
244
+
245
+ Before launching more reviewers, run one non-terminal convergence audit after
246
+ two consecutive Closure waves in one gate, ten total implementation review
247
+ passes, or 90 minutes when timing is available. One coordinated launch over all
248
+ affected lineages is one wave regardless of parallel pass count.
249
+
250
+ Persist one audit epoch with trigger, triggering pass/wave/revision,
251
+ completion, dispositions, selected sessions, and resume reason. It survives
252
+ resume. After a material repair/evidence change proves progress, rearm by
253
+ recording the current total pass count, current gate-local Closure-wave count,
254
+ and current timestamp as new baselines. The next audit opens only after a
255
+ post-rearm delta reaches two Closure waves in that gate, ten implementation
256
+ review passes, or 90 minutes; already-consumed counts or time cannot reopen it
257
+ immediately. Without progress do not rearm and use the existing stop rules.
258
+ Thresholds never approve, waive, downgrade, or block by themselves, and
259
+ distinct failure mechanics remain distinct even when they protect one
260
+ invariant.
261
+
262
+ Return `Approved` for the final settled revision only when protocol state is
263
+ `clear` and:
264
+
265
+ - every mandatory lens has independent coverage
266
+ - final validation and required cleanup/code-review gates ran
267
+
268
+ Map protocol `stopped` to `Blocked`. Implementation-specific blockers also
269
+ include:
270
+
271
+ - a defect reopens repeatedly and exposes a source-of-truth or repair-design
272
+ contradiction that root cannot resolve from current evidence
273
+ - an execution risk remains open without an explicit user decision to accept it
274
+
275
+ Do not mark implementation `Blocked` merely because review has run several
276
+ times. Resolve the repair, evidence, or decision problem first.
277
+
278
+ Map protocol `waived` to `Waived` and preserve skipped coverage and open risks.
279
+ It remains non-approval; target authority or downstream policy may still block
280
+ delivery.
281
+
282
+ ## Required Handoff
283
+
284
+ Use the protocol result envelope and add:
285
+
286
+ ```text
287
+ Implementation Review Profile: <simple | medium | high>
288
+ Review Outcome: <Approved | Blocked | Waived>
289
+ Authority Artifact: <approved spec path>
290
+ Implementation Checkpoint: <checkpoint or final>
291
+ ```
292
+
293
+ ## Contract Test Ledger
294
+
295
+ | Invariant | Risk It Prevents | First Test / Proof | Status |
296
+ | --- | --- | --- | --- |
297
+ | Review pass counts are audit metrics, while one durable Review Plan and Defect Ledger span the entire spec. | Each slice silently recreates a new loop or a repairable spec blocks on an arbitrary count. | Manual eval scenario 15 | planned |
298
+ | Intermediate checkpoints never trigger cleanup; final review is one settled profile-selected wave, and separate cleanup is exceptional. | Per-slice hygiene or size-driven cleanup adds latency and restarts review over unstable work. | Manual eval scenario 17 | planned |
299
+ | Audit epochs persist across resume and rearm only after material progress. | Thresholds repeatedly trigger audits or become terminal limits. | Manual eval scenario 12 | planned |
300
+
301
+ Keep these rows `planned` until the corresponding operator eval is run and its
302
+ result is saved.
@@ -0,0 +1,49 @@
1
+ # Review Gates
2
+
3
+ This file owns applicability only. Approved-spec execution follows
4
+ [`implementation-review-loop.md`](implementation-review-loop.md), which owns
5
+ checkpoint order, Full/Closure topology, durable state, and final coverage.
6
+ Direct work uses the gates below without manufacturing Module state.
7
+
8
+ ## Cleanup Review
9
+
10
+ Do not launch a separate `$cleanup-review` from size or risk classification
11
+ alone. The final `$code-review` spec/standards lens owns bounded cleanup for
12
+ simple, medium, large, and high-risk changes; high runs that lens in its own
13
+ parallel reviewer track.
14
+
15
+ Use separate `$cleanup-review` only when the user, approved source, or repo
16
+ policy names a concrete evidenced simplification risk that cannot fit the
17
+ bounded spec/standards lens. It runs before final `$code-review`; approved-spec
18
+ execution lets the shared Module schedule it.
19
+
20
+ Treat a change as large only when it contains several independently verifiable
21
+ runtime workflows or material cross-owner, cross-repo, release-sequencing, or
22
+ rollback coordination. File count, module count, or a broad mechanical diff is
23
+ not enough. Default a coherent feature with one behavior and one validation
24
+ path to medium even when it touches several files.
25
+
26
+ Cleanup is not a correctness review. Its skill owns the detailed lens,
27
+ confidence handling, repair integration, and output contract.
28
+
29
+ ## Final Code Review
30
+
31
+ Run final `$code-review` when implementation changes include any of:
32
+
33
+ - a medium/large feature or shared multi-module business behavior;
34
+ - API/DTO/schema, migration, persistence, auth, permission, payment, cache,
35
+ concurrency, background-job, or shared-state contracts;
36
+ - shared UI/navigation/middleware/core flows;
37
+ - runtime logic across three or more files when the change is behavioral rather
38
+ than mechanical and crosses one owner or validation seam.
39
+
40
+ Do not invoke it automatically for documentation/copy/comments, tests-only or
41
+ styling-only changes, formatting/renames, mechanical refactors, or isolated
42
+ one-file fixes with low regression risk.
43
+
44
+ `$code-review` owns reviewer topology, confidence, auto-fix, and output rules.
45
+ Its spec/standards reviewer owns the bounded cleanup lens for every profile; do
46
+ not launch `$cleanup-review` for the same settled diff without an explicit
47
+ concrete reason beyond size or risk labels.
48
+ For approved specs, required final coverage remains part of the shared
49
+ Implementation Review State.
@@ -0,0 +1,170 @@
1
+ # Review Protocol
2
+
3
+ This Module owns review mechanics shared by artifact and implementation review:
4
+ capsules, Full/Closure, reviewer lineage, the canonical Defect Ledger, no-progress,
5
+ stop/waiver semantics, and the result envelope. Target Modules own authority,
6
+ risk/profile selection, topology, durable orchestration, gate order, and final
7
+ outcome mapping. Callers select a target Module and never copy this protocol.
8
+
9
+ ## Interface
10
+
11
+ ```text
12
+ apply_review_protocol(
13
+ review_capsule,
14
+ reviewer_lineages
15
+ ) -> protocol_result
16
+ ```
17
+
18
+ `review_capsule` is the normalized Full/Closure capsule defined below. The result
19
+ contains `state: clear | repair-required | stopped | waived`, pass
20
+ mode/revision/lineage/session/lenses, mandatory coverage, ledger transitions, and any
21
+ stop or waiver record. The capsule's canonical Defect Ledger is the only ledger
22
+ input. `clear` requires assigned-scope coverage and every non-superseded blocker
23
+ or execution risk to be `verified`, except an explicitly `accepted-risk`
24
+ execution risk. A superseded record counts as resolved only when `superseded_by`
25
+ resolves through an acyclic chain of distinct same-ledger IDs to one
26
+ non-superseded canonical replacement; that terminal replacement controls
27
+ clearance. A missing, self-referential, or cyclic replacement chain prevents
28
+ `clear`. The target Module decides whether `clear` is enough for `Approved`. A
29
+ reviewer lineage binds one activation, profile/lenses, Full coverage, and defect
30
+ IDs; physical sessions may rotate without creating a new activation or Full.
31
+
32
+ ## Review Capsule
33
+
34
+ Every fresh reviewer receives a bounded self-contained capsule:
35
+
36
+ - target kind/path, pinned revision, profile, mode, and assigned lenses
37
+ - one review question and why existing valid coverage does not answer it
38
+ - goal, authority, approved scope, and relevant source references
39
+ - required Evidence Index entries, compact validation, and verification gaps
40
+ - canonical Defect Ledger
41
+ - for Closure: repair diff, affected contracts, and Revision Map
42
+ `changed target -> defect IDs -> affected invariants/contracts`
43
+
44
+ Exclude raw parent history, old targets, prior reviewer prose, repeated logs,
45
+ unrelated output, and broad inventories. Reviewers may inspect extra evidence
46
+ needed for a finding. Reuse unchanged evidence; invalidate only entries affected
47
+ by changed targets/contracts or stale external sources. External evidence keeps
48
+ its direct source and retrieval date.
49
+
50
+ Before launch, reuse clear coverage for the same target revision, question, and
51
+ lenses. Do not launch a reviewer whose question is already answered; a changed
52
+ target or a distinct artifact-versus-implementation question is new coverage.
53
+
54
+ ## Review Modes
55
+
56
+ **Full:** read the complete pinned target required by assigned lenses and return
57
+ all visible evidence-backed blockers/execution risks as one batch.
58
+
59
+ **Closure:** use only affected reviewer lineages to verify repaired IDs, the
60
+ repair diff, and contract fan-out. Do not re-audit unchanged areas. The first
61
+ Closure normally reuses the Full session. If it reopens a defect and another
62
+ repair follows, run the next Closure in a fresh session of the same lineage.
63
+ Before a Closure launch, rotate earlier after compaction/interruption or when
64
+ observable context usage reaches 40%. The fresh session receives the bounded
65
+ capsule, remains Closure, and preserves Full coverage and defect authority. One
66
+ coordinated launch over affected lineages is one Closure wave.
67
+
68
+ Closure admits a new defect only when it names the repair target/contract that
69
+ introduced or materially widened the trigger, or a high-confidence
70
+ `blocker`/`execution-risk` violates a source-required invariant in the assigned
71
+ lens. Each newly admitted blocker or execution risk starts as `status: open`
72
+ and `disposition: repair-now`, then moves `open -> fixed -> verified` only
73
+ through repair and independent affected-lens verification. Pre-existing
74
+ adjacent issues are improvements. Verification stays with an affected session
75
+ or an already planned covering Full reviewer. Start a new Full only when the
76
+ repair invalidated mandatory-lens coverage.
77
+
78
+ ## Canonical Defect Ledger
79
+
80
+ Root assigns stable IDs and deduplicates only when invariant and failure
81
+ mechanics both match. Wording changes do not create defects.
82
+
83
+ ```yaml
84
+ id: REVIEW-CONC-003
85
+ class: blocker | execution-risk | improvement
86
+ disposition: repair-now | planned-final-verification | follow-up-improvement | accepted-risk
87
+ status: open | fixed | verified | blocked | reopened | accepted-risk | superseded
88
+ severity: critical | high | medium | low
89
+ confidence: high | medium | low
90
+ invariant: "<observable rule>"
91
+ failure: "<concrete failure mechanics>"
92
+ evidence: ["<target/repo/source reference>"]
93
+ repair: "<smallest sufficient change>"
94
+ affected_targets: ["<path, section, contract, or lens>"]
95
+ introduced_in_review: 2
96
+ transition_history: []
97
+ reopen_count: 0
98
+ acceptance: null | { authority, reason, scope, target_revision }
99
+ superseded_by: null | "<canonical replacement defect ID>"
100
+ fixed_in_revision: null
101
+ verified_in_review: null
102
+ ```
103
+
104
+ Reviewers reuse supplied IDs; new candidates use `NEW-<LENS>-NN` until root
105
+ deduplicates them. Root may mark a duplicate `superseded` only after recording
106
+ its canonical replacement in `superseded_by`. The chain must be acyclic and end
107
+ at a distinct non-superseded record; a superseded record never hides the
108
+ terminal replacement's lifecycle. Disposition and lifecycle are separate.
109
+ `fixed` requires a later affected-lens Closure or planned Full to become
110
+ `verified`; tests alone do not independently verify review findings.
111
+
112
+ Improvements use `follow-up-improvement` and do not block. Only an execution
113
+ risk may become `accepted-risk`, after explicit authority, reason, scope, and
114
+ target revision are recorded. A blocker cannot be accepted or downgraded.
115
+ Target Modules decide where `planned-final-verification` is legal; it remains
116
+ open until the scheduled lens verifies it.
117
+
118
+ ## Repair, Stop, And Waiver
119
+
120
+ Root repairs compatible findings in one batch. Before another Closure, target
121
+ revision/strategy, relevant evidence, defect state, or source decision must
122
+ change materially. Otherwise root must change the repair, prove the finding
123
+ invalid, or surface the decision preventing convergence.
124
+
125
+ Pass counts and elapsed time are audit signals, never approval, waiver,
126
+ downgrade, or blocking conditions. Target Modules may define audit epochs while
127
+ preserving this rule.
128
+
129
+ A bounded poll timeout while the reviewer session remains live is non-terminal:
130
+ report that the reviewer has not completed within the polling window, not that
131
+ it is stuck, failed, or unavailable. Root may send at most one conclude request
132
+ for that reviewer turn. Further empty polls neither authorize another conclude
133
+ request nor replacement/cancellation. Replace or cancel only after the reviewer
134
+ session or transport explicitly reports `failed`, `lost`, or unavailable; a
135
+ late usable result from the original live session remains authoritative. Persist
136
+ session `live|timed-out` state and `conclude_requested_at` so resume preserves
137
+ the limit.
138
+
139
+ Return `stopped` without another review when repair needs a product/scope/owner
140
+ decision, mandatory evidence or reviewer is unavailable, no substantive repair
141
+ exists, or the same defect repeats without progress.
142
+
143
+ Return `waived` only after explicit user instruction. Record skipped coverage
144
+ and open defects. Waiver never means approval, verifies/accepts no defect, and
145
+ skips no separate downstream gate. The target Module maps it to `Waived` or
146
+ `Blocked` using its authority rules.
147
+
148
+ ## Result Envelope
149
+
150
+ ```text
151
+ Review Passes: <total; Full; Closure; fresh sessions when tracked>
152
+ Mandatory Coverage: <covered lenses or gaps>
153
+ Verified Defects: <IDs or None>
154
+ Accepted Risks: <IDs, authority, and reason or None>
155
+ Open Defects: <IDs or None>
156
+ ```
157
+
158
+ Target Modules add profile, mapped outcome, authority, Adapter verdict,
159
+ checkpoint, status, or performance fields without redefining common fields.
160
+
161
+ ## Contract Test Ledger
162
+
163
+ | Invariant | Risk It Prevents | First Test / Proof | Status |
164
+ | --- | --- | --- | --- |
165
+ | Full/Closure stay causal and affected-lens-only within reviewer lineage | Broad review restarts, stale context, or lost defect authority | Evals 11, 16 | green |
166
+ | Every launch answers an uncovered question on its pinned target | Duplicate reviewers repeat valid coverage without new evidence | Eval 23 | green |
167
+ | Both consumers use one defect schema and lifecycle | Artifact and implementation defects drift | Eval 12 | green |
168
+ | No-progress requires material change | Unchanged Closure repeats or numeric limits become terminal | Eval 12 | green |
169
+ | A live reviewer poll timeout stays non-terminal and permits at most one conclude request | Root falsely reports a hang, cancels useful work, or launches a duplicate reviewer | Eval 12 | green |
170
+ | Waiver skips coverage without accepting defects | Skipped review is reported as approval or risk acceptance | Eval 12 | green |