codex-orchestrator 2.0.3 → 2.0.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (161) hide show
  1. package/CHANGELOG.md +51 -427
  2. package/README.md +161 -37
  3. package/dist/src/index.d.ts +1 -1
  4. package/dist/src/index.d.ts.map +1 -1
  5. package/dist/src/v2/acceptance-proof.d.ts +5 -0
  6. package/dist/src/v2/acceptance-proof.d.ts.map +1 -1
  7. package/dist/src/v2/acceptance-proof.js +10 -2
  8. package/dist/src/v2/acceptance-proof.js.map +1 -1
  9. package/dist/src/v2/adapters/gh-issue-adapter.d.ts.map +1 -1
  10. package/dist/src/v2/adapters/gh-issue-adapter.js +6 -7
  11. package/dist/src/v2/adapters/gh-issue-adapter.js.map +1 -1
  12. package/dist/src/v2/adapters/gh-pull-request-adapter.d.ts +7 -1
  13. package/dist/src/v2/adapters/gh-pull-request-adapter.d.ts.map +1 -1
  14. package/dist/src/v2/adapters/gh-pull-request-adapter.js +288 -0
  15. package/dist/src/v2/adapters/gh-pull-request-adapter.js.map +1 -1
  16. package/dist/src/v2/adapters/pull-requests.d.ts +69 -0
  17. package/dist/src/v2/adapters/pull-requests.d.ts.map +1 -1
  18. package/dist/src/v2/adapters/pull-requests.js +48 -0
  19. package/dist/src/v2/adapters/pull-requests.js.map +1 -1
  20. package/dist/src/v2/adapters/worktree.d.ts +1 -0
  21. package/dist/src/v2/adapters/worktree.d.ts.map +1 -1
  22. package/dist/src/v2/adapters/worktree.js +10 -1
  23. package/dist/src/v2/adapters/worktree.js.map +1 -1
  24. package/dist/src/v2/cli-contract.d.ts +3 -3
  25. package/dist/src/v2/cli-contract.d.ts.map +1 -1
  26. package/dist/src/v2/cli-contract.js +1 -3
  27. package/dist/src/v2/cli-contract.js.map +1 -1
  28. package/dist/src/v2/cli.d.ts +33 -0
  29. package/dist/src/v2/cli.d.ts.map +1 -0
  30. package/dist/src/v2/{candidate-cli.js → cli.js} +55 -39
  31. package/dist/src/v2/cli.js.map +1 -0
  32. package/dist/src/v2/code-review-report.d.ts +1 -1
  33. package/dist/src/v2/code-review-report.d.ts.map +1 -1
  34. package/dist/src/v2/code-review-report.js +2 -2
  35. package/dist/src/v2/code-review-report.js.map +1 -1
  36. package/dist/src/v2/codex-process.d.ts.map +1 -1
  37. package/dist/src/v2/codex-process.js +12 -1
  38. package/dist/src/v2/codex-process.js.map +1 -1
  39. package/dist/src/v2/config.d.ts +2 -3
  40. package/dist/src/v2/config.d.ts.map +1 -1
  41. package/dist/src/v2/config.js +0 -3
  42. package/dist/src/v2/config.js.map +1 -1
  43. package/dist/src/v2/contained-report-operation.d.ts +2 -2
  44. package/dist/src/v2/contained-report-operation.d.ts.map +1 -1
  45. package/dist/src/v2/contained-report-operation.js +1 -1
  46. package/dist/src/v2/contained-report-operation.js.map +1 -1
  47. package/dist/src/v2/containment.d.ts +15 -2
  48. package/dist/src/v2/containment.d.ts.map +1 -1
  49. package/dist/src/v2/containment.js +43 -6
  50. package/dist/src/v2/containment.js.map +1 -1
  51. package/dist/src/v2/direct-delivery.d.ts +5 -10
  52. package/dist/src/v2/direct-delivery.d.ts.map +1 -1
  53. package/dist/src/v2/direct-delivery.js +32 -91
  54. package/dist/src/v2/direct-delivery.js.map +1 -1
  55. package/dist/src/v2/proof-report.d.ts.map +1 -1
  56. package/dist/src/v2/proof-report.js +55 -29
  57. package/dist/src/v2/proof-report.js.map +1 -1
  58. package/dist/src/v2/review-feedback-coordinator.d.ts +54 -0
  59. package/dist/src/v2/review-feedback-coordinator.d.ts.map +1 -0
  60. package/dist/src/v2/review-feedback-coordinator.js +245 -0
  61. package/dist/src/v2/review-feedback-coordinator.js.map +1 -0
  62. package/dist/src/v2/review-feedback.d.ts +127 -0
  63. package/dist/src/v2/review-feedback.d.ts.map +1 -0
  64. package/dist/src/v2/review-feedback.js +436 -0
  65. package/dist/src/v2/review-feedback.js.map +1 -0
  66. package/dist/src/v2/run-issue.d.ts +63 -9
  67. package/dist/src/v2/run-issue.d.ts.map +1 -1
  68. package/dist/src/v2/run-issue.js +798 -78
  69. package/dist/src/v2/run-issue.js.map +1 -1
  70. package/dist/src/v2/run-store.d.ts +49 -5
  71. package/dist/src/v2/run-store.d.ts.map +1 -1
  72. package/dist/src/v2/run-store.js +138 -44
  73. package/dist/src/v2/run-store.js.map +1 -1
  74. package/dist/src/v2/runtime.d.ts +40 -3
  75. package/dist/src/v2/runtime.d.ts.map +1 -1
  76. package/dist/src/v2/runtime.js +245 -52
  77. package/dist/src/v2/runtime.js.map +1 -1
  78. package/dist/src/v2/setup-cli.d.ts.map +1 -1
  79. package/dist/src/v2/setup-cli.js +4 -11
  80. package/dist/src/v2/setup-cli.js.map +1 -1
  81. package/dist/src/v2/setup-runtime.d.ts.map +1 -1
  82. package/dist/src/v2/setup-runtime.js +1 -61
  83. package/dist/src/v2/setup-runtime.js.map +1 -1
  84. package/dist/src/v2/setup-store.d.ts +0 -5
  85. package/dist/src/v2/setup-store.d.ts.map +1 -1
  86. package/dist/src/v2/setup-store.js +3 -106
  87. package/dist/src/v2/setup-store.js.map +1 -1
  88. package/dist/src/v2/setup.d.ts +6 -46
  89. package/dist/src/v2/setup.d.ts.map +1 -1
  90. package/dist/src/v2/setup.js +12 -294
  91. package/dist/src/v2/setup.js.map +1 -1
  92. package/dist/src/v2/workflow-assets.d.ts +19 -11
  93. package/dist/src/v2/workflow-assets.d.ts.map +1 -1
  94. package/dist/src/v2/workflow-assets.js +132 -40
  95. package/dist/src/v2/workflow-assets.js.map +1 -1
  96. package/docs/deep-dive.md +328 -56
  97. package/internal-workflow/docs/agents/bugfix-quality-gate.md +11 -0
  98. package/internal-workflow/docs/agents/coding-skill-routing.md +116 -196
  99. package/internal-workflow/docs/agents/contract-test-ledger.md +11 -1
  100. package/internal-workflow/docs/agents/review-gates.md +32 -39
  101. package/internal-workflow/docs/agents/review-protocol.md +75 -147
  102. package/internal-workflow/evals/coding-skill-evals.json +84 -0
  103. package/internal-workflow/manifest.json +1 -1
  104. package/internal-workflow/operations/acceptance-proof/SKILL.md +7 -1
  105. package/internal-workflow/operations/ambiguity-review/SKILL.md +2 -0
  106. package/internal-workflow/operations/code-review/SKILL.md +21 -1
  107. package/internal-workflow/operations/implementation/SKILL.md +22 -1
  108. package/internal-workflow/operations/spec-author/SKILL.md +10 -1
  109. package/internal-workflow/operations/spec-review/SKILL.md +10 -1
  110. package/internal-workflow/operations/triage/SKILL.md +10 -1
  111. package/internal-workflow/schemas/code-review-v1.json +1 -1
  112. package/internal-workflow/schemas/proof-report-v1.json +1 -1
  113. package/internal-workflow/skills/agent-auto/SKILL.md +6 -1
  114. package/internal-workflow/skills/code-debugger/SKILL.md +122 -0
  115. package/internal-workflow/skills/code-debugger/agents/openai.yaml +7 -0
  116. package/internal-workflow/skills/code-review/SKILL.md +51 -17
  117. package/internal-workflow/skills/code-review/references/cleanup-lens.md +52 -0
  118. package/internal-workflow/skills/implementation-spec-maker/SKILL.md +15 -6
  119. package/internal-workflow/skills/implementation-spec-maker/agents/openai.yaml +1 -1
  120. package/internal-workflow/skills/implementation-spec-maker/references/spec-template.md +2 -2
  121. package/internal-workflow/skills/implementation-spec-review/SKILL.md +108 -204
  122. package/internal-workflow/skills/implementation-spec-review/evals/evals.json +24 -0
  123. package/internal-workflow/skills/implementation-spec-review/references/review-loop.md +93 -0
  124. package/internal-workflow/skills/small-task-implementer/SKILL.md +15 -8
  125. package/internal-workflow/skills/spec-implementer/SKILL.md +101 -172
  126. package/internal-workflow/skills/spec-implementer/evals/evals.json +30 -0
  127. package/internal-workflow/skills/spec-implementer/references/review-loop.md +100 -0
  128. package/internal-workflow/skills/tdd/SKILL.md +20 -6
  129. package/internal-workflow/skills/tdd/agents/openai.yaml +2 -2
  130. package/internal-workflow/skills/tdd/evals/evals.json +18 -0
  131. package/internal-workflow/skills/tdd/mocking.md +3 -42
  132. package/internal-workflow/skills/tdd/refactoring.md +6 -8
  133. package/package.json +9 -6
  134. package/dist/src/v2/adapters/target-activity-fence.d.ts +0 -23
  135. package/dist/src/v2/adapters/target-activity-fence.d.ts.map +0 -1
  136. package/dist/src/v2/adapters/target-activity-fence.js +0 -249
  137. package/dist/src/v2/adapters/target-activity-fence.js.map +0 -1
  138. package/dist/src/v2/candidate-cli.d.ts +0 -26
  139. package/dist/src/v2/candidate-cli.d.ts.map +0 -1
  140. package/dist/src/v2/candidate-cli.js.map +0 -1
  141. package/dist/src/v2/legacy-cutover.d.ts +0 -52
  142. package/dist/src/v2/legacy-cutover.d.ts.map +0 -1
  143. package/dist/src/v2/legacy-cutover.js +0 -87
  144. package/dist/src/v2/legacy-cutover.js.map +0 -1
  145. package/internal-workflow/docs/agents/artifact-review-loop.md +0 -267
  146. package/internal-workflow/docs/agents/implementation-review-loop.md +0 -302
  147. package/internal-workflow/operations/cleanup-review/SKILL.md +0 -3
  148. package/internal-workflow/operations/spec-implementation/SKILL.md +0 -3
  149. package/internal-workflow/profiles/implementer_deep.toml +0 -9
  150. package/internal-workflow/profiles/researcher_standard.toml +0 -9
  151. package/internal-workflow/profiles/reviewer_fast.toml +0 -9
  152. package/internal-workflow/skills/cleanup-review/SKILL.md +0 -84
  153. package/internal-workflow/skills/cleanup-review/agents/openai.yaml +0 -6
  154. package/internal-workflow/skills/codebase-design/DEEPENING.md +0 -35
  155. package/internal-workflow/skills/codebase-design/DESIGN-IT-TWICE.md +0 -50
  156. package/internal-workflow/skills/codebase-design/SKILL.md +0 -82
  157. package/internal-workflow/skills/codebase-design/agents/openai.yaml +0 -6
  158. package/internal-workflow/skills/research/SKILL.md +0 -107
  159. package/internal-workflow/skills/research/agents/openai.yaml +0 -6
  160. package/internal-workflow/skills/ui-evidence-proof/SKILL.md +0 -123
  161. package/internal-workflow/skills/ui-evidence-proof/agents/openai.yaml +0 -6
@@ -5,207 +5,111 @@ description: "Review compact or full implementation specs for deterministic exec
5
5
 
6
6
  # Implementation Spec Review
7
7
 
8
- Review an implementation spec before execution begins. The spec is an execution contract for a downstream coding agent. Your job is to decide whether it can be executed safely without guessing, not whether the product idea is good.
9
-
10
- Use `../../docs/agents/confidence-rubric.md` when classifying blockers, execution risks, and optional improvements. High-confidence blockers need direct evidence from the spec, repo, issue, or trusted contract. Medium-confidence risks must name the one unresolved assumption. Low-confidence concerns are questions or verification gaps, not blockers.
11
-
12
- Use `../../docs/agents/contract-test-ledger.md` when reviewing behavior-changing specs with contract risk.
13
-
14
- This skill reviews three independent classifications from
15
- `implementation-spec-maker`: document shape (`spec_mode`), delivery shape
16
- (`implementation_size`), and consequence/uncertainty (`review_profile`). It
17
- also checks the declared `expected_repositories`. Never infer one dimension
18
- from another.
19
-
20
- It supports both document shapes:
21
-
22
- - **Compact specs:** dense single-agent specs whose exact scope, risk controls, and proof fit clearly. Compact does not mean small; a coherent medium/large or high-risk implementation may stay compact when ownership, sequencing, and validation remain deterministic. They do not need source-of-truth tables, file matrices, multi-agent contracts, or long halt sections unless a concrete ambiguity calls for them.
23
- - **Full specs:** lean contracts used when compact mode cannot express concrete coordination, contract, ownership, sequencing, or validation ambiguity safely.
24
-
25
- Do not punish a compact spec for omitting full-mode ceremony. Do not punish a full spec for using short risk-control bullets instead of large tables when the ownership and safety rules are still unambiguous. Do reject any spec, compact or full, that requires invention, hides risk, or cannot prove the intended behavior.
26
-
27
- When used inside the spec-making loop, return blockers and author questions clearly so the parent agent can relay them. Do not contact the user directly. Do not rewrite the spec unless explicitly asked.
28
- When already running as the assigned `reviewer_fast`, `reviewer_standard`, or
29
- `reviewer_deep` child, review inline and never spawn a grandchild. Otherwise,
30
- root must launch the reviewer selected by the shared review profile and must not
31
- self-review inline.
32
-
33
- ## Review Posture
34
-
35
- - Be strict about determinism and safety.
36
- - Be proportional about format and ceremony.
37
- - Prefer concrete defects with repair instructions over broad commentary, but in Full Mode return every visible evidence-backed blocker and execution risk in the assigned lenses as one batch.
38
- - Treat false precision as a first-class defect: exact-looking claims must be grounded in source material, repo evidence, or trusted external sources.
39
- - Separate hard blockers from execution risks and optional improvements.
40
- - If the spec can be fixed by one sentence, name that sentence-level fix instead of expanding the spec.
41
-
42
- ### Scope Conservation
43
-
44
- A review repair must not broaden approved product or operational scope. Prefer
45
- deleting or narrowing an unsafe proposal before adding a mechanism. Feature
46
- flags, telemetry systems, dashboards, rollout machinery, compatibility paths,
47
- and generic fallbacks are scope expansion unless the source requires them or a
48
- concrete evidenced failure path makes them necessary. Keep optional
49
- improvements optional; do not convert them into mandatory spec work.
50
-
51
- ## Artifact Review Protocol
52
-
53
- Act as the implementation-spec Adapter for
54
- `../../docs/agents/artifact-review-loop.md`. The prompt assigns mode and lenses.
55
- Return actual coverage, reuse supplied IDs, and label new candidates
56
- `NEW-<LENS>-NN`; do not implement another review flow or lifecycle here.
57
-
58
- When invoked directly outside a maker Module, default to `Full` mode, cover all
59
- mandatory spec lenses, use no ledger unless one is supplied, and return only the
60
- single Adapter verdict. Do not claim a Module outcome or closure state.
61
-
62
- ## Size-Aware Rules
63
-
64
- ### Compact Specs
65
-
66
- Approve a compact spec when it has:
67
-
68
- - exact enough targets, commands, preconditions, and validation for the current task
69
- - numbered phases or a clear single phase
70
- - a simple progress/reconciliation rule
71
- - explicit blockers or `None`
72
- - observable behavior proof
73
- - no unresolved placeholders, pseudo-paths, or alternative commands
74
-
75
- Do not require these unless the task risk demands them:
76
-
77
- - source-of-truth map
78
- - file modification matrix
79
- - long halt checklist
80
- - defect closure section
81
- - multi-agent handoff contract
82
- - mandatory dedicated review at every phase
83
-
84
- ### Lean Full Specs
85
-
86
- Treat these as signals to check whether compact mode leaves a concrete
87
- ambiguity; none selects full mode by itself:
88
-
89
- - multi-agent execution
90
- - persistence, migrations, schemas, DTOs, API contracts, or external contracts
91
- - auth, secrets, permissions, payments, caching, concurrency, background jobs, shared state, or destructive operations
92
- - changes across several runtime surfaces
93
- - one ticket with multi-agent coordination requirements
94
- - revision of an existing checklist with completed history
95
-
96
- For these specs, the expected shape is:
97
-
98
- - a short `Risk Controls` section naming only applicable ownership, safety, contract, concurrency/state, and forbidden-scope rules
99
- - phase steps with exact targets and validation
100
- - `Write Scope Summary` only when phases alone do not make the write set obvious, or when there is multi-agent work, broad runtime change, generated artifacts, or 5+ runtime files
101
- - task-specific `Halt Conditions` only when the compact stop rule is insufficient
102
- - `Integrator Coordination Contract` only for multi-agent execution
103
-
104
- Missing ownership, write-scope, validation, safety, or handoff details are defects. Missing tables are not defects when the lean sections are unambiguous.
105
-
106
- ## Mandatory Review Lenses
107
-
108
- Across a Module full-review wave, cover all of these lenses according to the
109
- policy assignment, scaled to artifact risk; each reviewer owns only its
110
- assigned primary lenses. A standalone direct review covers all lenses:
111
-
112
- - **Determinism:** exact paths, symbols, commands, payloads, fixtures, and target behavior where execution depends on them.
113
- - **Evidence:** exact-looking claims are supported by source material, repo context, docs, issues, or external contract proof.
114
- - **Preconditions:** required services, env vars, fixtures, data state, feature flags, and prerequisites are explicit or intentionally `None`.
115
- - **Sequencing:** phases are ordered safely and have exit checks.
116
- - **Scope:** approved scope, out of scope, protected paths, and rejected approaches are clear enough to prevent drift.
117
- - **Contract Test Ledger:** contract-heavy behavior changes map ordering, precedence, threading, runtime contract, retry/idempotency, determinism, evidence, partial failure, and cardinality risks to first tests/proofs.
118
- - **Review Checkpoints and Focus:** high-risk specs use an early `$code-review` checkpoint only when the first risky slice becomes stable before later work. Otherwise the spec assigns concrete `Review Focus` lenses, targeted recipes, and bug classes to the final parallel review wave.
119
- - **Final Handoff Requirements:** medium/high-risk specs require a compact final response packet covering contract implemented, risky checkpoints, invariants proved, review findings/fixes, validation, skipped checks, residual risks, and files by role.
120
- - **Minimum solution and reuse:** the spec states a direct `Minimum Solution`, records `Added Complexity: None` or ties every added mechanism to a concrete invariant/failure, and removes anything that passes the deletion challenge. Prefer the fewest necessary concepts, owners, states, and integration points; line or file count is not decisive.
121
- - **Deep-module fit:** using the `$codebase-design` lens, new Modules or Seams pass the deletion test, avoid one-adapter abstraction, and test through the Module Interface.
122
- - **Risk Controls:** full specs include only applicable risk controls, and each control is specific enough to guide execution.
123
- - **Ownership:** source-of-truth ownership is explicit when behavior/data can drift across layers. A short `Source of Truth` bullet is enough when there is only one material owner.
124
- - **Validation:** checks prove observable behavior, not just compilation.
125
- - **Safety:** auth, secrets, credentials, destructive operations, persistence, concurrency, retries, and shared state have explicit constraints when touched.
126
- - **Multi-agent handoff:** if multi-agent, write scopes are disjoint and one integrator owns merge order and final reconciliation.
127
- - **Revision integrity:** if revising, still-valid completed items are preserved and stale completed items are reopened with a note.
128
- - **Completion clarity:** another agent could know when to stop, what passed, and what remains blocked.
129
-
130
- ## What To Reject Immediately
131
-
132
- - The spec asks the executor to guess file names, symbol names, DTOs, schema details, API contracts, fixtures, or behavior.
133
- - The spec contains unresolved placeholders, pseudo-paths, example rows, bracket instructions, or alternative commands presented as executable.
134
- - The validation cannot prove the intended behavior.
135
- - The spec changes a contract-heavy behavior but has no Contract Test Ledger, or the ledger lists invariants without a first RED test/proof or a concrete blocked reason.
136
- - A high-risk spec neither defines a stable early checkpoint nor assigns the risky slice's concrete review focus to the final parallel wave.
137
- - A medium/high-risk spec has no final handoff requirement, leaving the user to manually reconstruct contract proof, review status, skipped checks, or residual risk from the diff.
138
- - Code changes are planned, the repo has an architecture check, and the spec omits it without a reason.
139
- - Exact-looking paths, commands, or symbols are not grounded in evidence and would force the executor to trust invented precision.
140
- - A multi-agent topology has overlapping write scopes, unclear integration ownership, or no merge/handoff contract.
141
- - A full spec touches a real safety/contract/state risk but has no applicable `Risk Controls` entry.
142
- - The spec says or implies the executor should continue despite a mismatch instead of stopping.
143
- - Security-sensitive or destructive work lacks explicit safe sources, guards, or stop-before-damage constraints.
144
- - The spec requires an added mechanism outside approved scope, or safe execution would depend on inventing that mechanism's need or contract.
145
-
146
- ## Common Defects
147
-
148
- - Vague instructions like “update logic”, “handle edge cases”, “refactor if needed”, or “reuse existing code where possible” without exact targets.
149
- - Validation that checks only lint/build and misses the behavior changed by the spec.
150
- - A contract-heavy spec covers only a happy path and omits ordering, precedence, retry/idempotency, persistence/evidence, serialization, deterministic ordering, or cardinality invariants that are material to the touched flow.
151
- - A high-risk spec defers all review to the final diff even though the first risky state/contract slice becomes stable, is independently reviewable, and will not be invalidated by lower-risk work.
152
- - A medium/high-risk spec updates checklists and review gates but does not say what proof summary the executor must give the user at completion.
153
- - Required env vars, fixtures, payloads, services, or app state are absent.
154
- - Acceptance criteria are subjective, non-observable, or not tied to proof.
155
- - Ticket specs lose issue-only requirements such as `Implementation preparation`, `External contracts`, `Verification`, `Blocked by`, live prerequisites, or rejected approaches, or absorb sibling-ticket scope.
156
- - Evidence sections repeat a file inventory instead of proving determinism.
157
- - `Minimum Solution` is a slogan rather than a direct path, `Added Complexity` is absent or vague, or a smaller evidence-backed path satisfies the same approved behavior, invariants, and proof.
158
- - Full specs add large tables where 2-4 exact `Risk Controls` bullets would be clearer.
159
- - Full specs include generic halt checklists instead of task-specific stop conditions.
160
- - Full specs spread one rule across multiple files without a declared owner.
161
- - Compact specs expand into a large document without added safety value.
162
-
163
- ## Defect Taxonomy
164
-
165
- - **Blocker:** The spec is unsafe or impossible to execute as written.
166
- - **Execution Risk:** The spec is executable but likely to cause drift, rework, or inconsistent implementation.
167
- - **Improvement:** The spec is usable, and the suggestion would materially sharpen it.
168
-
169
- Report blockers first. Mention improvements only when they matter.
170
-
171
- ## Decision Rules
172
-
173
- - **Approved:** Deterministic, bounded, proportionate, and executable without guesswork.
174
- - **Needs Work:** Directionally usable but has ambiguities, missing proof, weak validation, or scope/control gaps.
175
- - **Rejected:** Unsafe to execute because it depends on invention, broad interpretation, overlapping ownership, missing validation, or missing safety controls.
176
-
177
- Scores:
178
-
179
- - `0`: missing or unsafe
180
- - `1`: partially specified or weakly proven
181
- - `2`: explicit and well-grounded
182
-
183
- ## Output Format
184
-
185
- Always answer in Russian, keeping technical terms in English where appropriate. Use this exact structure:
186
-
187
- 1. `Вердикт: Approved / Needs Work / Rejected` plus one sentence with the main reason.
188
- 2. `Режим и покрытие: Full / Closure` with assigned lenses and evidence actually checked.
189
- 3. `Оценка` with short scores `Determinism / Evidence / Validation / Safety` on a 0-2 scale.
190
- 4. `Что уже исполнимо` with 2-4 bullets about what is concrete and safe.
191
- 5. `Критические дефекты спецификации` with concrete blockers, ambiguity points, and failure mechanics. Quote the exact vague phrase, missing step, or unsafe instruction when justifying a defect. If there are no blockers, say `Нет`.
192
- 6. `Defect Records` with the supplied stable ID or `NEW-<LENS>-NN`, class, confidence, invariant, failure, evidence, repair, affected sections, and status. If there are no defects, say `Нет`.
193
- 7. `Что исправить перед исполнением` with exact changes needed in the spec. If nothing is needed, say `Ничего`.
194
- 8. `Жесткие уточняющие вопросы` with 3-5 specific questions only if the spec cannot become deterministic without answers. If none, say `Нет`.
195
-
196
- Keep the output short, severe, and execution-oriented.
197
-
198
- ## Anti-Overengineering Heuristic
199
-
200
- - If a compact direct flow is enough, flag unnecessary full-mode ceremony as an improvement or execution risk.
201
- - Run the whole-solution deletion challenge: if a mechanism can be removed while preserving every approved behavior, material invariant, and proof, require its removal. Do not remove a necessary mechanism merely because it adds a file, type, schema object, or boundary.
202
- - If a lean full spec gives exact risk controls without tables, do not ask for tables unless prose leaves ambiguity.
203
- - If a full spec has enough phase-level targets, do not ask for `Write Scope Summary` unless the write set or ownership is hard to audit.
204
- - If a full spec introduces an indirect flow where a direct one satisfies all constraints, treat that as a defect.
205
- - If a new abstraction exists only for cleanliness or future flexibility, return `Needs Work` and require its removal. Treat it as a blocker only when execution would be unsafe, broaden approved scope, or require invention.
206
- - Treat missing or unsupported `Minimum Solution` / `Added Complexity` evidence as `Needs Work`; escalate to `Rejected` only when the added mechanism is unsafe, unapproved scope, or requires invention.
207
- - If the review can make the spec safer by deleting ceremony rather than adding it, say so.
208
-
209
- ## Tone
210
-
211
- Be direct, strict, and operational. No fluff, no architecture theater.
8
+ Decide whether a saved implementation spec can be executed safely without
9
+ guessing. Review execution quality, not the product idea. Do not rewrite the
10
+ spec unless explicitly asked.
11
+
12
+ Read:
13
+
14
+ - `references/review-loop.md` when called by `$implementation-spec-maker`;
15
+ - `../../docs/agents/confidence-rubric.md` for defect confidence;
16
+ - `../../docs/agents/contract-test-ledger.md` only when the spec changes a
17
+ material behavior contract.
18
+
19
+ ## Independent Dimensions
20
+
21
+ Keep these classifications independent:
22
+
23
+ - `spec_mode: compact | full` — document/coordination density;
24
+ - `implementation_size: small | medium | large` — delivery shape;
25
+ - `review_profile: simple | medium | high` — consequence and uncertainty;
26
+ - `expected_repositories` — approved repository count.
27
+
28
+ Compact may describe broad or high-risk work when ownership, sequencing, and
29
+ proof remain deterministic. Full is justified only when concrete coordination,
30
+ contract, safety, ownership, or validation ambiguity cannot fit clearly in the
31
+ compact form. Never request full-mode tables or ceremony merely from size or
32
+ risk labels.
33
+
34
+ ## Adapter Contract
35
+
36
+ When called by the maker, use the mode and lenses supplied by
37
+ `references/review-loop.md`, reuse supplied defect IDs, and return actual
38
+ coverage. A reviewer child executes this Adapter inline and never spawns a
39
+ grandchild. If root receives a direct review request, it launches the
40
+ profile-selected reviewer instead of self-reviewing.
41
+
42
+ A standalone reviewer performs one bounded Full over all applicable lenses and
43
+ returns only `Approved | Needs Work | Rejected`; it does not invent owner state
44
+ or claim Closure.
45
+
46
+ ## Review Lenses
47
+
48
+ Scale depth to the profile and inspect only applicable lenses:
49
+
50
+ - **Determinism and evidence:** execution-critical paths, symbols, commands,
51
+ contracts, fixtures, and claims are confirmed rather than invented.
52
+ - **Scope and minimum solution:** the spec preserves approved scope, uses
53
+ existing owners/seams, and ties every added mechanism to a requirement or
54
+ concrete failure path.
55
+ - **Sequencing and ownership:** phases are safe, sources of truth are explicit
56
+ where drift is possible, and multi-agent write scopes are disjoint.
57
+ - **Validation:** each behavior has an observable proof; contract-risk work maps
58
+ each material invariant to its first failing test or exact blocked proof.
59
+ - **Preconditions and stop conditions:** required services, data, env, fixtures,
60
+ and destructive/sensitive constraints are explicit when applicable.
61
+ - **Review focus:** ordinary work relies on one final review; only an explicit
62
+ stable high-risk slice gets an intermediate checkpoint.
63
+ - **Revision integrity:** current content matches its authority and preserves
64
+ still-valid completed work.
65
+ - **Completion:** another agent can tell what to do, what proves success, when
66
+ to stop, and what remains blocked.
67
+
68
+ ## Proportional Expectations
69
+
70
+ Approve a compact spec when targets, ordered work, observable proof, and stop
71
+ conditions are exact enough for the task. Do not require source-of-truth tables,
72
+ file matrices, long halt lists, multi-agent contracts, or defect sections when
73
+ no concrete ambiguity needs them.
74
+
75
+ A lean full spec normally adds only applicable `Risk Controls`, exact phase
76
+ targets/proof, and—when needed—write-scope or integrator coordination. Missing
77
+ ownership, validation, safety, or handoff detail is a defect; missing formatting
78
+ ceremony is not.
79
+
80
+ Prefer deleting or narrowing an unsafe proposal before adding flags, telemetry,
81
+ fallbacks, compatibility paths, or rollout machinery. Optional improvements
82
+ remain optional unless source authority approves them.
83
+
84
+ ## Defects And Decision
85
+
86
+ - **Blocker:** unsafe or impossible to execute as written.
87
+ - **Execution risk:** executable but likely to drift or require rework.
88
+ - **Improvement:** useful but not required for safe execution.
89
+
90
+ Reject exact-looking but ungrounded paths/contracts, unresolved placeholders or
91
+ alternative commands, validation that cannot prove the intended behavior,
92
+ overlapping multi-agent ownership, missing material safety constraints, or any
93
+ step that requires invention.
94
+
95
+ Use:
96
+
97
+ - `Approved` when the current spec is deterministic, bounded, proportional, and
98
+ executable without guessing;
99
+ - `Needs Work` for repairable ambiguity, weak proof, or excess ceremony;
100
+ - `Rejected` when execution would be unsafe or depend on invented decisions.
101
+
102
+ ## Output
103
+
104
+ Answer in Russian and keep technical terms in English. Return:
105
+
106
+ 1. `Вердикт` and one-sentence reason.
107
+ 2. `Режим и покрытие` with Full/Closure and actual lenses.
108
+ 3. Short `Determinism / Evidence / Validation / Safety` scores from 0 to 2.
109
+ 4. Evidence-backed defects first, with supplied ID or `NEW-<LENS>-NN`, class,
110
+ confidence, failure, evidence, smallest repair, and affected section.
111
+ 5. Exact changes needed before execution, or `Ничего`.
112
+ 6. Only genuinely blocking questions, or `Нет`.
113
+
114
+ Do not repeat the spec, propose broad redesign, or turn optional cleanup into a
115
+ mandatory gate.
@@ -0,0 +1,24 @@
1
+ {
2
+ "schema_version": 1,
3
+ "skill": "implementation-spec-review",
4
+ "cases": [
5
+ {
6
+ "id": "artifact-profile-by-consequence",
7
+ "prompt": "Review one broad but reversible spec and one narrow spec with irreversible data impact and unclear recovery ownership.",
8
+ "expected": ["broad reversible may remain medium", "narrow dangerous uncertain spec is high"],
9
+ "forbidden": ["classify from file count or spec mode"]
10
+ },
11
+ {
12
+ "id": "artifact-scope-conservation",
13
+ "prompt": "A review can repair the spec either by deleting an unnecessary mechanism or by adding flags, telemetry, and fallback infrastructure.",
14
+ "expected": ["prefer the smallest scope-preserving repair"],
15
+ "forbidden": ["add unapproved operational machinery"]
16
+ },
17
+ {
18
+ "id": "artifact-approval-invalidation",
19
+ "prompt": "An approved spec receives a substantive execution change after review.",
20
+ "expected": ["invalidate approval for the changed revision", "review only invalidated coverage"],
21
+ "forbidden": ["execute under stale approval", "restart unrelated coverage"]
22
+ }
23
+ ]
24
+ }
@@ -0,0 +1,93 @@
1
+ # Implementation Spec Review Loop
2
+
3
+ This reference owns review orchestration for specs created by
4
+ `implementation-spec-maker`. Read it when the maker requests artifact review.
5
+ The reviewer skill remains the Adapter; the shared review mechanics live in
6
+ `../../../docs/agents/review-protocol.md`.
7
+
8
+ ## Contract
9
+
10
+ Input:
11
+
12
+ - saved spec and pinned revision;
13
+ - source authority and approved decisions;
14
+ - evidence needed to verify execution claims;
15
+ - optional user-raised review profile.
16
+
17
+ Output:
18
+
19
+ - `outcome: Approved | Blocked | Waived`;
20
+ - `adapter_verdict: Approved | Needs Work | Rejected | Not run`;
21
+ - `review_profile: simple | medium | high` and evidence-backed reasons;
22
+ - mandatory-lens coverage and unresolved defects.
23
+
24
+ The Adapter returns only its verdict. The root maps preflight, convergence, and
25
+ waiver state to the artifact outcome.
26
+
27
+ ## Preflight And Profile
28
+
29
+ Before launching a reviewer, confirm source authority, approved scope, current
30
+ spec revision, and mandatory external evidence. Save a useful blocked spec when
31
+ a product or contract decision is missing; do not launch review to discover a
32
+ known authority gap.
33
+
34
+ `medium` is the default. Use:
35
+
36
+ - `simple` for one narrow owner with direct proof and no material uncertainty;
37
+ - `medium` for all ordinary specs, including multi-file, API, persistence, or
38
+ stateful work with clear ownership and bounded proof;
39
+ - `high` only when a sensitive mechanism has both a material failure
40
+ consequence and an uncertainty amplifier such as unclear ownership,
41
+ cross-trust effects, non-local recovery, or an unproven external contract.
42
+
43
+ File count and implementation size never select `high`. The user may raise but
44
+ not lower an evidence-backed profile.
45
+
46
+ ## Scope And Capsule
47
+
48
+ Review the smallest approved solution. Risk may strengthen proof but does not
49
+ authorize flags, telemetry, compatibility paths, generic fallbacks, or rollout
50
+ machinery unless the source or a concrete failure requires them.
51
+
52
+ Give each reviewer a bounded capsule containing the current spec, authority,
53
+ approved scope, evidence, review question, assigned lenses, and current defect
54
+ records. For Closure also include the repaired sections and affected contracts.
55
+ Do not pass raw parent history or unrelated inventories.
56
+
57
+ ## Topology
58
+
59
+ - `simple`: one `reviewer_fast`, one bounded Full.
60
+ - `medium`: one `reviewer_standard`, one bounded Full.
61
+ - `high`: two parallel `reviewer_deep` sessions with disjoint primary lenses:
62
+ Architecture/Execution and Failure/Contracts.
63
+
64
+ Root launches and aggregates reviewers. A reviewer child runs the
65
+ `implementation-spec-review` Adapter inline and never spawns another reviewer.
66
+ Reuse valid coverage for the same revision and question.
67
+
68
+ After one consolidated repair, coordinator verification is enough for ordinary
69
+ medium/low findings. Use shared-protocol Closure only for critical/high defects,
70
+ protected trust/data/concurrency/shared-contract impact, or invalidated
71
+ mandatory coverage. A substantive rewrite gets a new Full only when it
72
+ invalidates existing mandatory lenses.
73
+
74
+ ## Approval
75
+
76
+ Return `Approved` only when the current saved revision matches source authority,
77
+ mandatory lenses are covered, and every blocking defect is verified. Any
78
+ substantive edit invalidates approval; lifecycle metadata alone does not.
79
+
80
+ Return `Blocked` when authority/evidence is missing, repair needs a product or
81
+ ownership decision, no substantive repair exists, or shared no-progress rules
82
+ apply. Return `Waived` only after explicit user instruction and keep skipped
83
+ coverage visible; an open blocker still maps the artifact to `Blocked`.
84
+
85
+ Map outcomes to spec status:
86
+
87
+ - `Approved` -> `ready`;
88
+ - `Blocked` -> `blocked`;
89
+ - eligible `Waived` -> `ready` with visible waiver metadata.
90
+
91
+ Report profile, outcome, Adapter verdict, mandatory coverage, verified/open
92
+ defects, and skipped checks. Do not report counters or session history for a
93
+ normal one-review flow.
@@ -16,15 +16,21 @@ Proceed only when all are true:
16
16
  - There is a narrow validation path: targeted test, lint/typecheck, build check, UI proof, or direct command.
17
17
  - The task does not require a new plan, PRD, issue breakdown, implementation spec, migration, rollout, or multi-agent orchestration.
18
18
 
19
- Escalate instead of implementing when the task touches:
20
-
21
- - state transitions, queues, retries, idempotency, background jobs, persistence, migrations, schemas, DTO/API contracts, auth, permissions, payments, caching, or shared cross-module behavior;
22
- - multi-service, multi-repo, multi-agent, production/live-data, or external-contract work;
23
- - unclear product intent, ambiguous scope, no credible validation path, or likely broad refactoring.
19
+ Escalate out of the tiny-task route when the work has more than one coherent
20
+ behavior, a broad ownership boundary, material rollback/recovery risk, unclear
21
+ product intent, no credible affected validation, or genuine multi-agent/live
22
+ coordination. Statefulness alone does not require escalation into planning or
23
+ orchestration; a clear API, persistence, cache, queue, or DTO change normally
24
+ becomes direct medium implementation.
24
25
 
25
26
  Escalation rule:
26
27
 
27
- - Escalate into the single canonical delivery flow: optional `$grilling` for unresolved product decisions, then `$spec-to-tickets` for the reviewed Approval Packet, then `$tickets-orchestrator` for approved ticket delivery.
28
+ - For clear authority and one coherent outcome, escalate to direct medium root
29
+ implementation under `$tdd`, affected validation, and one final review when
30
+ `review-gates.md` applies.
31
+ - Use optional `$grilling`, then `$spec-to-tickets` and `$tickets-orchestrator`,
32
+ only for unresolved product decisions or a real approved ticket graph,
33
+ delivery dependency, or explicit orchestration request.
28
34
  - For one risky behavior or technical contract, prefer one approved ticket and mark `compact spec` or `standard spec` only when the ticket plus repository evidence cannot remove execution ambiguity.
29
35
  - For several tickets sharing one unresolved contract or validation path, make the contract-defining ticket block its consumers; merge tickets that cannot be specified or verified independently instead of creating a wave-level implementation spec.
30
36
  - Escalate if the bug requires Bugfix Quality Gate analysis across multiple paths, states, async events, persistence, auth, cache, retries, workers, or contracts.
@@ -58,7 +64,8 @@ Validation:
58
64
  - Do not run full CI unless local policy or the changed surface makes it necessary.
59
65
 
60
66
  5. Stop and escalate if implementation reveals hidden risk.
61
- - Examples: shared contract drift, duplicate source of truth, missing test seam, broad file spread, concurrency, persistence, or product ambiguity.
67
+ - Examples: shared contract drift, duplicate source of truth, missing test
68
+ seam, material concurrency/recovery uncertainty, or product ambiguity.
62
69
  - Leave a short explanation of what was discovered and which heavier flow should take over.
63
70
 
64
71
  ## Output
@@ -90,7 +97,7 @@ Reason:
90
97
  - ...
91
98
 
92
99
  Recommended flow:
93
- - Canonical ticket delivery / direct small task
100
+ - Direct medium implementation / canonical ticket delivery
94
101
 
95
102
  Evidence:
96
103
  - ...