codex-orchestrator 2.0.3 → 2.0.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +28 -427
- package/README.md +135 -37
- package/dist/src/index.d.ts +1 -1
- package/dist/src/index.d.ts.map +1 -1
- package/dist/src/v2/adapters/gh-issue-adapter.d.ts.map +1 -1
- package/dist/src/v2/adapters/gh-issue-adapter.js +6 -7
- package/dist/src/v2/adapters/gh-issue-adapter.js.map +1 -1
- package/dist/src/v2/cli-contract.d.ts +3 -3
- package/dist/src/v2/cli-contract.d.ts.map +1 -1
- package/dist/src/v2/cli-contract.js +1 -3
- package/dist/src/v2/cli-contract.js.map +1 -1
- package/dist/src/v2/cli.d.ts +24 -0
- package/dist/src/v2/cli.d.ts.map +1 -0
- package/dist/src/v2/{candidate-cli.js → cli.js} +18 -29
- package/dist/src/v2/cli.js.map +1 -0
- package/dist/src/v2/code-review-report.d.ts +1 -1
- package/dist/src/v2/code-review-report.d.ts.map +1 -1
- package/dist/src/v2/code-review-report.js +2 -2
- package/dist/src/v2/code-review-report.js.map +1 -1
- package/dist/src/v2/codex-process.d.ts.map +1 -1
- package/dist/src/v2/codex-process.js +12 -1
- package/dist/src/v2/codex-process.js.map +1 -1
- package/dist/src/v2/config.d.ts +2 -2
- package/dist/src/v2/config.d.ts.map +1 -1
- package/dist/src/v2/config.js.map +1 -1
- package/dist/src/v2/contained-report-operation.d.ts +2 -2
- package/dist/src/v2/contained-report-operation.d.ts.map +1 -1
- package/dist/src/v2/contained-report-operation.js +1 -1
- package/dist/src/v2/contained-report-operation.js.map +1 -1
- package/dist/src/v2/containment.d.ts +4 -0
- package/dist/src/v2/containment.d.ts.map +1 -1
- package/dist/src/v2/containment.js +9 -0
- package/dist/src/v2/containment.js.map +1 -1
- package/dist/src/v2/direct-delivery.d.ts +5 -10
- package/dist/src/v2/direct-delivery.d.ts.map +1 -1
- package/dist/src/v2/direct-delivery.js +25 -90
- package/dist/src/v2/direct-delivery.js.map +1 -1
- package/dist/src/v2/proof-report.d.ts.map +1 -1
- package/dist/src/v2/proof-report.js +55 -29
- package/dist/src/v2/proof-report.js.map +1 -1
- package/dist/src/v2/run-issue.d.ts +6 -9
- package/dist/src/v2/run-issue.d.ts.map +1 -1
- package/dist/src/v2/run-issue.js +98 -43
- package/dist/src/v2/run-issue.js.map +1 -1
- package/dist/src/v2/run-store.d.ts +4 -4
- package/dist/src/v2/run-store.d.ts.map +1 -1
- package/dist/src/v2/run-store.js +25 -40
- package/dist/src/v2/run-store.js.map +1 -1
- package/dist/src/v2/runtime.d.ts +3 -3
- package/dist/src/v2/runtime.d.ts.map +1 -1
- package/dist/src/v2/runtime.js +125 -47
- package/dist/src/v2/runtime.js.map +1 -1
- package/dist/src/v2/setup-cli.d.ts.map +1 -1
- package/dist/src/v2/setup-cli.js +4 -11
- package/dist/src/v2/setup-cli.js.map +1 -1
- package/dist/src/v2/setup-runtime.d.ts.map +1 -1
- package/dist/src/v2/setup-runtime.js +1 -61
- package/dist/src/v2/setup-runtime.js.map +1 -1
- package/dist/src/v2/setup-store.d.ts +0 -5
- package/dist/src/v2/setup-store.d.ts.map +1 -1
- package/dist/src/v2/setup-store.js +3 -106
- package/dist/src/v2/setup-store.js.map +1 -1
- package/dist/src/v2/setup.d.ts +6 -46
- package/dist/src/v2/setup.d.ts.map +1 -1
- package/dist/src/v2/setup.js +11 -293
- package/dist/src/v2/setup.js.map +1 -1
- package/dist/src/v2/workflow-assets.d.ts +19 -11
- package/dist/src/v2/workflow-assets.d.ts.map +1 -1
- package/dist/src/v2/workflow-assets.js +132 -40
- package/dist/src/v2/workflow-assets.js.map +1 -1
- package/docs/deep-dive.md +272 -56
- package/internal-workflow/docs/agents/bugfix-quality-gate.md +11 -0
- package/internal-workflow/docs/agents/coding-skill-routing.md +116 -196
- package/internal-workflow/docs/agents/review-gates.md +32 -39
- package/internal-workflow/docs/agents/review-protocol.md +75 -147
- package/internal-workflow/evals/coding-skill-evals.json +66 -0
- package/internal-workflow/manifest.json +1 -1
- package/internal-workflow/operations/acceptance-proof/SKILL.md +7 -1
- package/internal-workflow/operations/ambiguity-review/SKILL.md +2 -0
- package/internal-workflow/operations/code-review/SKILL.md +21 -1
- package/internal-workflow/operations/implementation/SKILL.md +22 -1
- package/internal-workflow/operations/spec-author/SKILL.md +10 -1
- package/internal-workflow/operations/spec-review/SKILL.md +10 -1
- package/internal-workflow/operations/triage/SKILL.md +10 -1
- package/internal-workflow/schemas/code-review-v1.json +1 -1
- package/internal-workflow/schemas/proof-report-v1.json +1 -1
- package/internal-workflow/skills/agent-auto/SKILL.md +6 -1
- package/internal-workflow/skills/code-debugger/SKILL.md +122 -0
- package/internal-workflow/skills/code-debugger/agents/openai.yaml +7 -0
- package/internal-workflow/skills/code-review/SKILL.md +33 -11
- package/internal-workflow/skills/code-review/references/cleanup-lens.md +52 -0
- package/internal-workflow/skills/implementation-spec-maker/SKILL.md +15 -6
- package/internal-workflow/skills/implementation-spec-maker/references/spec-template.md +2 -2
- package/internal-workflow/skills/implementation-spec-review/SKILL.md +108 -204
- package/internal-workflow/skills/implementation-spec-review/evals/evals.json +24 -0
- package/internal-workflow/skills/implementation-spec-review/references/review-loop.md +93 -0
- package/internal-workflow/skills/small-task-implementer/SKILL.md +15 -8
- package/internal-workflow/skills/spec-implementer/SKILL.md +101 -172
- package/internal-workflow/skills/spec-implementer/evals/evals.json +30 -0
- package/internal-workflow/skills/spec-implementer/references/review-loop.md +94 -0
- package/internal-workflow/skills/tdd/SKILL.md +15 -2
- package/internal-workflow/skills/tdd/agents/openai.yaml +2 -2
- package/package.json +9 -6
- package/dist/src/v2/adapters/target-activity-fence.d.ts +0 -23
- package/dist/src/v2/adapters/target-activity-fence.d.ts.map +0 -1
- package/dist/src/v2/adapters/target-activity-fence.js +0 -249
- package/dist/src/v2/adapters/target-activity-fence.js.map +0 -1
- package/dist/src/v2/candidate-cli.d.ts +0 -26
- package/dist/src/v2/candidate-cli.d.ts.map +0 -1
- package/dist/src/v2/candidate-cli.js.map +0 -1
- package/dist/src/v2/legacy-cutover.d.ts +0 -52
- package/dist/src/v2/legacy-cutover.d.ts.map +0 -1
- package/dist/src/v2/legacy-cutover.js +0 -87
- package/dist/src/v2/legacy-cutover.js.map +0 -1
- package/internal-workflow/docs/agents/artifact-review-loop.md +0 -267
- package/internal-workflow/docs/agents/implementation-review-loop.md +0 -302
- package/internal-workflow/operations/cleanup-review/SKILL.md +0 -3
- package/internal-workflow/operations/spec-implementation/SKILL.md +0 -3
- package/internal-workflow/profiles/implementer_deep.toml +0 -9
- package/internal-workflow/profiles/researcher_standard.toml +0 -9
- package/internal-workflow/profiles/reviewer_fast.toml +0 -9
- package/internal-workflow/skills/cleanup-review/SKILL.md +0 -84
- package/internal-workflow/skills/cleanup-review/agents/openai.yaml +0 -6
- package/internal-workflow/skills/codebase-design/DEEPENING.md +0 -35
- package/internal-workflow/skills/codebase-design/DESIGN-IT-TWICE.md +0 -50
- package/internal-workflow/skills/codebase-design/SKILL.md +0 -82
- package/internal-workflow/skills/codebase-design/agents/openai.yaml +0 -6
- package/internal-workflow/skills/research/SKILL.md +0 -107
- package/internal-workflow/skills/research/agents/openai.yaml +0 -6
- package/internal-workflow/skills/ui-evidence-proof/SKILL.md +0 -123
- package/internal-workflow/skills/ui-evidence-proof/agents/openai.yaml +0 -6
|
@@ -5,207 +5,111 @@ description: "Review compact or full implementation specs for deterministic exec
|
|
|
5
5
|
|
|
6
6
|
# Implementation Spec Review
|
|
7
7
|
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
-
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
-
|
|
91
|
-
|
|
92
|
-
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
-
|
|
100
|
-
- `
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
- **Scope:** approved scope, out of scope, protected paths, and rejected approaches are clear enough to prevent drift.
|
|
117
|
-
- **Contract Test Ledger:** contract-heavy behavior changes map ordering, precedence, threading, runtime contract, retry/idempotency, determinism, evidence, partial failure, and cardinality risks to first tests/proofs.
|
|
118
|
-
- **Review Checkpoints and Focus:** high-risk specs use an early `$code-review` checkpoint only when the first risky slice becomes stable before later work. Otherwise the spec assigns concrete `Review Focus` lenses, targeted recipes, and bug classes to the final parallel review wave.
|
|
119
|
-
- **Final Handoff Requirements:** medium/high-risk specs require a compact final response packet covering contract implemented, risky checkpoints, invariants proved, review findings/fixes, validation, skipped checks, residual risks, and files by role.
|
|
120
|
-
- **Minimum solution and reuse:** the spec states a direct `Minimum Solution`, records `Added Complexity: None` or ties every added mechanism to a concrete invariant/failure, and removes anything that passes the deletion challenge. Prefer the fewest necessary concepts, owners, states, and integration points; line or file count is not decisive.
|
|
121
|
-
- **Deep-module fit:** using the `$codebase-design` lens, new Modules or Seams pass the deletion test, avoid one-adapter abstraction, and test through the Module Interface.
|
|
122
|
-
- **Risk Controls:** full specs include only applicable risk controls, and each control is specific enough to guide execution.
|
|
123
|
-
- **Ownership:** source-of-truth ownership is explicit when behavior/data can drift across layers. A short `Source of Truth` bullet is enough when there is only one material owner.
|
|
124
|
-
- **Validation:** checks prove observable behavior, not just compilation.
|
|
125
|
-
- **Safety:** auth, secrets, credentials, destructive operations, persistence, concurrency, retries, and shared state have explicit constraints when touched.
|
|
126
|
-
- **Multi-agent handoff:** if multi-agent, write scopes are disjoint and one integrator owns merge order and final reconciliation.
|
|
127
|
-
- **Revision integrity:** if revising, still-valid completed items are preserved and stale completed items are reopened with a note.
|
|
128
|
-
- **Completion clarity:** another agent could know when to stop, what passed, and what remains blocked.
|
|
129
|
-
|
|
130
|
-
## What To Reject Immediately
|
|
131
|
-
|
|
132
|
-
- The spec asks the executor to guess file names, symbol names, DTOs, schema details, API contracts, fixtures, or behavior.
|
|
133
|
-
- The spec contains unresolved placeholders, pseudo-paths, example rows, bracket instructions, or alternative commands presented as executable.
|
|
134
|
-
- The validation cannot prove the intended behavior.
|
|
135
|
-
- The spec changes a contract-heavy behavior but has no Contract Test Ledger, or the ledger lists invariants without a first RED test/proof or a concrete blocked reason.
|
|
136
|
-
- A high-risk spec neither defines a stable early checkpoint nor assigns the risky slice's concrete review focus to the final parallel wave.
|
|
137
|
-
- A medium/high-risk spec has no final handoff requirement, leaving the user to manually reconstruct contract proof, review status, skipped checks, or residual risk from the diff.
|
|
138
|
-
- Code changes are planned, the repo has an architecture check, and the spec omits it without a reason.
|
|
139
|
-
- Exact-looking paths, commands, or symbols are not grounded in evidence and would force the executor to trust invented precision.
|
|
140
|
-
- A multi-agent topology has overlapping write scopes, unclear integration ownership, or no merge/handoff contract.
|
|
141
|
-
- A full spec touches a real safety/contract/state risk but has no applicable `Risk Controls` entry.
|
|
142
|
-
- The spec says or implies the executor should continue despite a mismatch instead of stopping.
|
|
143
|
-
- Security-sensitive or destructive work lacks explicit safe sources, guards, or stop-before-damage constraints.
|
|
144
|
-
- The spec requires an added mechanism outside approved scope, or safe execution would depend on inventing that mechanism's need or contract.
|
|
145
|
-
|
|
146
|
-
## Common Defects
|
|
147
|
-
|
|
148
|
-
- Vague instructions like “update logic”, “handle edge cases”, “refactor if needed”, or “reuse existing code where possible” without exact targets.
|
|
149
|
-
- Validation that checks only lint/build and misses the behavior changed by the spec.
|
|
150
|
-
- A contract-heavy spec covers only a happy path and omits ordering, precedence, retry/idempotency, persistence/evidence, serialization, deterministic ordering, or cardinality invariants that are material to the touched flow.
|
|
151
|
-
- A high-risk spec defers all review to the final diff even though the first risky state/contract slice becomes stable, is independently reviewable, and will not be invalidated by lower-risk work.
|
|
152
|
-
- A medium/high-risk spec updates checklists and review gates but does not say what proof summary the executor must give the user at completion.
|
|
153
|
-
- Required env vars, fixtures, payloads, services, or app state are absent.
|
|
154
|
-
- Acceptance criteria are subjective, non-observable, or not tied to proof.
|
|
155
|
-
- Ticket specs lose issue-only requirements such as `Implementation preparation`, `External contracts`, `Verification`, `Blocked by`, live prerequisites, or rejected approaches, or absorb sibling-ticket scope.
|
|
156
|
-
- Evidence sections repeat a file inventory instead of proving determinism.
|
|
157
|
-
- `Minimum Solution` is a slogan rather than a direct path, `Added Complexity` is absent or vague, or a smaller evidence-backed path satisfies the same approved behavior, invariants, and proof.
|
|
158
|
-
- Full specs add large tables where 2-4 exact `Risk Controls` bullets would be clearer.
|
|
159
|
-
- Full specs include generic halt checklists instead of task-specific stop conditions.
|
|
160
|
-
- Full specs spread one rule across multiple files without a declared owner.
|
|
161
|
-
- Compact specs expand into a large document without added safety value.
|
|
162
|
-
|
|
163
|
-
## Defect Taxonomy
|
|
164
|
-
|
|
165
|
-
- **Blocker:** The spec is unsafe or impossible to execute as written.
|
|
166
|
-
- **Execution Risk:** The spec is executable but likely to cause drift, rework, or inconsistent implementation.
|
|
167
|
-
- **Improvement:** The spec is usable, and the suggestion would materially sharpen it.
|
|
168
|
-
|
|
169
|
-
Report blockers first. Mention improvements only when they matter.
|
|
170
|
-
|
|
171
|
-
## Decision Rules
|
|
172
|
-
|
|
173
|
-
- **Approved:** Deterministic, bounded, proportionate, and executable without guesswork.
|
|
174
|
-
- **Needs Work:** Directionally usable but has ambiguities, missing proof, weak validation, or scope/control gaps.
|
|
175
|
-
- **Rejected:** Unsafe to execute because it depends on invention, broad interpretation, overlapping ownership, missing validation, or missing safety controls.
|
|
176
|
-
|
|
177
|
-
Scores:
|
|
178
|
-
|
|
179
|
-
- `0`: missing or unsafe
|
|
180
|
-
- `1`: partially specified or weakly proven
|
|
181
|
-
- `2`: explicit and well-grounded
|
|
182
|
-
|
|
183
|
-
## Output Format
|
|
184
|
-
|
|
185
|
-
Always answer in Russian, keeping technical terms in English where appropriate. Use this exact structure:
|
|
186
|
-
|
|
187
|
-
1. `Вердикт: Approved / Needs Work / Rejected` plus one sentence with the main reason.
|
|
188
|
-
2. `Режим и покрытие: Full / Closure` with assigned lenses and evidence actually checked.
|
|
189
|
-
3. `Оценка` with short scores `Determinism / Evidence / Validation / Safety` on a 0-2 scale.
|
|
190
|
-
4. `Что уже исполнимо` with 2-4 bullets about what is concrete and safe.
|
|
191
|
-
5. `Критические дефекты спецификации` with concrete blockers, ambiguity points, and failure mechanics. Quote the exact vague phrase, missing step, or unsafe instruction when justifying a defect. If there are no blockers, say `Нет`.
|
|
192
|
-
6. `Defect Records` with the supplied stable ID or `NEW-<LENS>-NN`, class, confidence, invariant, failure, evidence, repair, affected sections, and status. If there are no defects, say `Нет`.
|
|
193
|
-
7. `Что исправить перед исполнением` with exact changes needed in the spec. If nothing is needed, say `Ничего`.
|
|
194
|
-
8. `Жесткие уточняющие вопросы` with 3-5 specific questions only if the spec cannot become deterministic without answers. If none, say `Нет`.
|
|
195
|
-
|
|
196
|
-
Keep the output short, severe, and execution-oriented.
|
|
197
|
-
|
|
198
|
-
## Anti-Overengineering Heuristic
|
|
199
|
-
|
|
200
|
-
- If a compact direct flow is enough, flag unnecessary full-mode ceremony as an improvement or execution risk.
|
|
201
|
-
- Run the whole-solution deletion challenge: if a mechanism can be removed while preserving every approved behavior, material invariant, and proof, require its removal. Do not remove a necessary mechanism merely because it adds a file, type, schema object, or boundary.
|
|
202
|
-
- If a lean full spec gives exact risk controls without tables, do not ask for tables unless prose leaves ambiguity.
|
|
203
|
-
- If a full spec has enough phase-level targets, do not ask for `Write Scope Summary` unless the write set or ownership is hard to audit.
|
|
204
|
-
- If a full spec introduces an indirect flow where a direct one satisfies all constraints, treat that as a defect.
|
|
205
|
-
- If a new abstraction exists only for cleanliness or future flexibility, return `Needs Work` and require its removal. Treat it as a blocker only when execution would be unsafe, broaden approved scope, or require invention.
|
|
206
|
-
- Treat missing or unsupported `Minimum Solution` / `Added Complexity` evidence as `Needs Work`; escalate to `Rejected` only when the added mechanism is unsafe, unapproved scope, or requires invention.
|
|
207
|
-
- If the review can make the spec safer by deleting ceremony rather than adding it, say so.
|
|
208
|
-
|
|
209
|
-
## Tone
|
|
210
|
-
|
|
211
|
-
Be direct, strict, and operational. No fluff, no architecture theater.
|
|
8
|
+
Decide whether a saved implementation spec can be executed safely without
|
|
9
|
+
guessing. Review execution quality, not the product idea. Do not rewrite the
|
|
10
|
+
spec unless explicitly asked.
|
|
11
|
+
|
|
12
|
+
Read:
|
|
13
|
+
|
|
14
|
+
- `references/review-loop.md` when called by `$implementation-spec-maker`;
|
|
15
|
+
- `../../docs/agents/confidence-rubric.md` for defect confidence;
|
|
16
|
+
- `../../docs/agents/contract-test-ledger.md` only when the spec changes a
|
|
17
|
+
material behavior contract.
|
|
18
|
+
|
|
19
|
+
## Independent Dimensions
|
|
20
|
+
|
|
21
|
+
Keep these classifications independent:
|
|
22
|
+
|
|
23
|
+
- `spec_mode: compact | full` — document/coordination density;
|
|
24
|
+
- `implementation_size: small | medium | large` — delivery shape;
|
|
25
|
+
- `review_profile: simple | medium | high` — consequence and uncertainty;
|
|
26
|
+
- `expected_repositories` — approved repository count.
|
|
27
|
+
|
|
28
|
+
Compact may describe broad or high-risk work when ownership, sequencing, and
|
|
29
|
+
proof remain deterministic. Full is justified only when concrete coordination,
|
|
30
|
+
contract, safety, ownership, or validation ambiguity cannot fit clearly in the
|
|
31
|
+
compact form. Never request full-mode tables or ceremony merely from size or
|
|
32
|
+
risk labels.
|
|
33
|
+
|
|
34
|
+
## Adapter Contract
|
|
35
|
+
|
|
36
|
+
When called by the maker, use the mode and lenses supplied by
|
|
37
|
+
`references/review-loop.md`, reuse supplied defect IDs, and return actual
|
|
38
|
+
coverage. A reviewer child executes this Adapter inline and never spawns a
|
|
39
|
+
grandchild. If root receives a direct review request, it launches the
|
|
40
|
+
profile-selected reviewer instead of self-reviewing.
|
|
41
|
+
|
|
42
|
+
A standalone reviewer performs one bounded Full over all applicable lenses and
|
|
43
|
+
returns only `Approved | Needs Work | Rejected`; it does not invent owner state
|
|
44
|
+
or claim Closure.
|
|
45
|
+
|
|
46
|
+
## Review Lenses
|
|
47
|
+
|
|
48
|
+
Scale depth to the profile and inspect only applicable lenses:
|
|
49
|
+
|
|
50
|
+
- **Determinism and evidence:** execution-critical paths, symbols, commands,
|
|
51
|
+
contracts, fixtures, and claims are confirmed rather than invented.
|
|
52
|
+
- **Scope and minimum solution:** the spec preserves approved scope, uses
|
|
53
|
+
existing owners/seams, and ties every added mechanism to a requirement or
|
|
54
|
+
concrete failure path.
|
|
55
|
+
- **Sequencing and ownership:** phases are safe, sources of truth are explicit
|
|
56
|
+
where drift is possible, and multi-agent write scopes are disjoint.
|
|
57
|
+
- **Validation:** each behavior has an observable proof; contract-risk work maps
|
|
58
|
+
each material invariant to its first failing test or exact blocked proof.
|
|
59
|
+
- **Preconditions and stop conditions:** required services, data, env, fixtures,
|
|
60
|
+
and destructive/sensitive constraints are explicit when applicable.
|
|
61
|
+
- **Review focus:** ordinary work relies on one final review; only an explicit
|
|
62
|
+
stable high-risk slice gets an intermediate checkpoint.
|
|
63
|
+
- **Revision integrity:** current content matches its authority and preserves
|
|
64
|
+
still-valid completed work.
|
|
65
|
+
- **Completion:** another agent can tell what to do, what proves success, when
|
|
66
|
+
to stop, and what remains blocked.
|
|
67
|
+
|
|
68
|
+
## Proportional Expectations
|
|
69
|
+
|
|
70
|
+
Approve a compact spec when targets, ordered work, observable proof, and stop
|
|
71
|
+
conditions are exact enough for the task. Do not require source-of-truth tables,
|
|
72
|
+
file matrices, long halt lists, multi-agent contracts, or defect sections when
|
|
73
|
+
no concrete ambiguity needs them.
|
|
74
|
+
|
|
75
|
+
A lean full spec normally adds only applicable `Risk Controls`, exact phase
|
|
76
|
+
targets/proof, and—when needed—write-scope or integrator coordination. Missing
|
|
77
|
+
ownership, validation, safety, or handoff detail is a defect; missing formatting
|
|
78
|
+
ceremony is not.
|
|
79
|
+
|
|
80
|
+
Prefer deleting or narrowing an unsafe proposal before adding flags, telemetry,
|
|
81
|
+
fallbacks, compatibility paths, or rollout machinery. Optional improvements
|
|
82
|
+
remain optional unless source authority approves them.
|
|
83
|
+
|
|
84
|
+
## Defects And Decision
|
|
85
|
+
|
|
86
|
+
- **Blocker:** unsafe or impossible to execute as written.
|
|
87
|
+
- **Execution risk:** executable but likely to drift or require rework.
|
|
88
|
+
- **Improvement:** useful but not required for safe execution.
|
|
89
|
+
|
|
90
|
+
Reject exact-looking but ungrounded paths/contracts, unresolved placeholders or
|
|
91
|
+
alternative commands, validation that cannot prove the intended behavior,
|
|
92
|
+
overlapping multi-agent ownership, missing material safety constraints, or any
|
|
93
|
+
step that requires invention.
|
|
94
|
+
|
|
95
|
+
Use:
|
|
96
|
+
|
|
97
|
+
- `Approved` when the current spec is deterministic, bounded, proportional, and
|
|
98
|
+
executable without guessing;
|
|
99
|
+
- `Needs Work` for repairable ambiguity, weak proof, or excess ceremony;
|
|
100
|
+
- `Rejected` when execution would be unsafe or depend on invented decisions.
|
|
101
|
+
|
|
102
|
+
## Output
|
|
103
|
+
|
|
104
|
+
Answer in Russian and keep technical terms in English. Return:
|
|
105
|
+
|
|
106
|
+
1. `Вердикт` and one-sentence reason.
|
|
107
|
+
2. `Режим и покрытие` with Full/Closure and actual lenses.
|
|
108
|
+
3. Short `Determinism / Evidence / Validation / Safety` scores from 0 to 2.
|
|
109
|
+
4. Evidence-backed defects first, with supplied ID or `NEW-<LENS>-NN`, class,
|
|
110
|
+
confidence, failure, evidence, smallest repair, and affected section.
|
|
111
|
+
5. Exact changes needed before execution, or `Ничего`.
|
|
112
|
+
6. Only genuinely blocking questions, or `Нет`.
|
|
113
|
+
|
|
114
|
+
Do not repeat the spec, propose broad redesign, or turn optional cleanup into a
|
|
115
|
+
mandatory gate.
|
|
@@ -0,0 +1,24 @@
|
|
|
1
|
+
{
|
|
2
|
+
"schema_version": 1,
|
|
3
|
+
"skill": "implementation-spec-review",
|
|
4
|
+
"cases": [
|
|
5
|
+
{
|
|
6
|
+
"id": "artifact-profile-by-consequence",
|
|
7
|
+
"prompt": "Review one broad but reversible spec and one narrow spec with irreversible data impact and unclear recovery ownership.",
|
|
8
|
+
"expected": ["broad reversible may remain medium", "narrow dangerous uncertain spec is high"],
|
|
9
|
+
"forbidden": ["classify from file count or spec mode"]
|
|
10
|
+
},
|
|
11
|
+
{
|
|
12
|
+
"id": "artifact-scope-conservation",
|
|
13
|
+
"prompt": "A review can repair the spec either by deleting an unnecessary mechanism or by adding flags, telemetry, and fallback infrastructure.",
|
|
14
|
+
"expected": ["prefer the smallest scope-preserving repair"],
|
|
15
|
+
"forbidden": ["add unapproved operational machinery"]
|
|
16
|
+
},
|
|
17
|
+
{
|
|
18
|
+
"id": "artifact-approval-invalidation",
|
|
19
|
+
"prompt": "An approved spec receives a substantive execution change after review.",
|
|
20
|
+
"expected": ["invalidate approval for the changed revision", "review only invalidated coverage"],
|
|
21
|
+
"forbidden": ["execute under stale approval", "restart unrelated coverage"]
|
|
22
|
+
}
|
|
23
|
+
]
|
|
24
|
+
}
|
|
@@ -0,0 +1,93 @@
|
|
|
1
|
+
# Implementation Spec Review Loop
|
|
2
|
+
|
|
3
|
+
This reference owns review orchestration for specs created by
|
|
4
|
+
`implementation-spec-maker`. Read it when the maker requests artifact review.
|
|
5
|
+
The reviewer skill remains the Adapter; the shared review mechanics live in
|
|
6
|
+
`../../../docs/agents/review-protocol.md`.
|
|
7
|
+
|
|
8
|
+
## Contract
|
|
9
|
+
|
|
10
|
+
Input:
|
|
11
|
+
|
|
12
|
+
- saved spec and pinned revision;
|
|
13
|
+
- source authority and approved decisions;
|
|
14
|
+
- evidence needed to verify execution claims;
|
|
15
|
+
- optional user-raised review profile.
|
|
16
|
+
|
|
17
|
+
Output:
|
|
18
|
+
|
|
19
|
+
- `outcome: Approved | Blocked | Waived`;
|
|
20
|
+
- `adapter_verdict: Approved | Needs Work | Rejected | Not run`;
|
|
21
|
+
- `review_profile: simple | medium | high` and evidence-backed reasons;
|
|
22
|
+
- mandatory-lens coverage and unresolved defects.
|
|
23
|
+
|
|
24
|
+
The Adapter returns only its verdict. The root maps preflight, convergence, and
|
|
25
|
+
waiver state to the artifact outcome.
|
|
26
|
+
|
|
27
|
+
## Preflight And Profile
|
|
28
|
+
|
|
29
|
+
Before launching a reviewer, confirm source authority, approved scope, current
|
|
30
|
+
spec revision, and mandatory external evidence. Save a useful blocked spec when
|
|
31
|
+
a product or contract decision is missing; do not launch review to discover a
|
|
32
|
+
known authority gap.
|
|
33
|
+
|
|
34
|
+
`medium` is the default. Use:
|
|
35
|
+
|
|
36
|
+
- `simple` for one narrow owner with direct proof and no material uncertainty;
|
|
37
|
+
- `medium` for all ordinary specs, including multi-file, API, persistence, or
|
|
38
|
+
stateful work with clear ownership and bounded proof;
|
|
39
|
+
- `high` only when a sensitive mechanism has both a material failure
|
|
40
|
+
consequence and an uncertainty amplifier such as unclear ownership,
|
|
41
|
+
cross-trust effects, non-local recovery, or an unproven external contract.
|
|
42
|
+
|
|
43
|
+
File count and implementation size never select `high`. The user may raise but
|
|
44
|
+
not lower an evidence-backed profile.
|
|
45
|
+
|
|
46
|
+
## Scope And Capsule
|
|
47
|
+
|
|
48
|
+
Review the smallest approved solution. Risk may strengthen proof but does not
|
|
49
|
+
authorize flags, telemetry, compatibility paths, generic fallbacks, or rollout
|
|
50
|
+
machinery unless the source or a concrete failure requires them.
|
|
51
|
+
|
|
52
|
+
Give each reviewer a bounded capsule containing the current spec, authority,
|
|
53
|
+
approved scope, evidence, review question, assigned lenses, and current defect
|
|
54
|
+
records. For Closure also include the repaired sections and affected contracts.
|
|
55
|
+
Do not pass raw parent history or unrelated inventories.
|
|
56
|
+
|
|
57
|
+
## Topology
|
|
58
|
+
|
|
59
|
+
- `simple`: one `reviewer_fast`, one bounded Full.
|
|
60
|
+
- `medium`: one `reviewer_standard`, one bounded Full.
|
|
61
|
+
- `high`: two parallel `reviewer_deep` sessions with disjoint primary lenses:
|
|
62
|
+
Architecture/Execution and Failure/Contracts.
|
|
63
|
+
|
|
64
|
+
Root launches and aggregates reviewers. A reviewer child runs the
|
|
65
|
+
`implementation-spec-review` Adapter inline and never spawns another reviewer.
|
|
66
|
+
Reuse valid coverage for the same revision and question.
|
|
67
|
+
|
|
68
|
+
After one consolidated repair, coordinator verification is enough for ordinary
|
|
69
|
+
medium/low findings. Use shared-protocol Closure only for critical/high defects,
|
|
70
|
+
protected trust/data/concurrency/shared-contract impact, or invalidated
|
|
71
|
+
mandatory coverage. A substantive rewrite gets a new Full only when it
|
|
72
|
+
invalidates existing mandatory lenses.
|
|
73
|
+
|
|
74
|
+
## Approval
|
|
75
|
+
|
|
76
|
+
Return `Approved` only when the current saved revision matches source authority,
|
|
77
|
+
mandatory lenses are covered, and every blocking defect is verified. Any
|
|
78
|
+
substantive edit invalidates approval; lifecycle metadata alone does not.
|
|
79
|
+
|
|
80
|
+
Return `Blocked` when authority/evidence is missing, repair needs a product or
|
|
81
|
+
ownership decision, no substantive repair exists, or shared no-progress rules
|
|
82
|
+
apply. Return `Waived` only after explicit user instruction and keep skipped
|
|
83
|
+
coverage visible; an open blocker still maps the artifact to `Blocked`.
|
|
84
|
+
|
|
85
|
+
Map outcomes to spec status:
|
|
86
|
+
|
|
87
|
+
- `Approved` -> `ready`;
|
|
88
|
+
- `Blocked` -> `blocked`;
|
|
89
|
+
- eligible `Waived` -> `ready` with visible waiver metadata.
|
|
90
|
+
|
|
91
|
+
Report profile, outcome, Adapter verdict, mandatory coverage, verified/open
|
|
92
|
+
defects, and skipped checks. Do not report counters or session history for a
|
|
93
|
+
normal one-review flow.
|
|
@@ -16,15 +16,21 @@ Proceed only when all are true:
|
|
|
16
16
|
- There is a narrow validation path: targeted test, lint/typecheck, build check, UI proof, or direct command.
|
|
17
17
|
- The task does not require a new plan, PRD, issue breakdown, implementation spec, migration, rollout, or multi-agent orchestration.
|
|
18
18
|
|
|
19
|
-
Escalate
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
19
|
+
Escalate out of the tiny-task route when the work has more than one coherent
|
|
20
|
+
behavior, a broad ownership boundary, material rollback/recovery risk, unclear
|
|
21
|
+
product intent, no credible affected validation, or genuine multi-agent/live
|
|
22
|
+
coordination. Statefulness alone does not require escalation into planning or
|
|
23
|
+
orchestration; a clear API, persistence, cache, queue, or DTO change normally
|
|
24
|
+
becomes direct medium implementation.
|
|
24
25
|
|
|
25
26
|
Escalation rule:
|
|
26
27
|
|
|
27
|
-
-
|
|
28
|
+
- For clear authority and one coherent outcome, escalate to direct medium root
|
|
29
|
+
implementation under `$tdd`, affected validation, and one final review when
|
|
30
|
+
`review-gates.md` applies.
|
|
31
|
+
- Use optional `$grilling`, then `$spec-to-tickets` and `$tickets-orchestrator`,
|
|
32
|
+
only for unresolved product decisions or a real approved ticket graph,
|
|
33
|
+
delivery dependency, or explicit orchestration request.
|
|
28
34
|
- For one risky behavior or technical contract, prefer one approved ticket and mark `compact spec` or `standard spec` only when the ticket plus repository evidence cannot remove execution ambiguity.
|
|
29
35
|
- For several tickets sharing one unresolved contract or validation path, make the contract-defining ticket block its consumers; merge tickets that cannot be specified or verified independently instead of creating a wave-level implementation spec.
|
|
30
36
|
- Escalate if the bug requires Bugfix Quality Gate analysis across multiple paths, states, async events, persistence, auth, cache, retries, workers, or contracts.
|
|
@@ -58,7 +64,8 @@ Validation:
|
|
|
58
64
|
- Do not run full CI unless local policy or the changed surface makes it necessary.
|
|
59
65
|
|
|
60
66
|
5. Stop and escalate if implementation reveals hidden risk.
|
|
61
|
-
- Examples: shared contract drift, duplicate source of truth, missing test
|
|
67
|
+
- Examples: shared contract drift, duplicate source of truth, missing test
|
|
68
|
+
seam, material concurrency/recovery uncertainty, or product ambiguity.
|
|
62
69
|
- Leave a short explanation of what was discovered and which heavier flow should take over.
|
|
63
70
|
|
|
64
71
|
## Output
|
|
@@ -90,7 +97,7 @@ Reason:
|
|
|
90
97
|
- ...
|
|
91
98
|
|
|
92
99
|
Recommended flow:
|
|
93
|
-
-
|
|
100
|
+
- Direct medium implementation / canonical ticket delivery
|
|
94
101
|
|
|
95
102
|
Evidence:
|
|
96
103
|
- ...
|