codex-orchestrator 2.0.11 → 2.0.12
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +15 -0
- package/README.md +25 -51
- package/dist/src/index.d.ts +2 -8
- package/dist/src/index.d.ts.map +1 -1
- package/dist/src/index.js +1 -4
- package/dist/src/index.js.map +1 -1
- package/dist/src/v2/acceptance-proof.d.ts +46 -31
- package/dist/src/v2/acceptance-proof.d.ts.map +1 -1
- package/dist/src/v2/acceptance-proof.js +157 -195
- package/dist/src/v2/acceptance-proof.js.map +1 -1
- package/dist/src/v2/active-attempt.d.ts +94 -0
- package/dist/src/v2/active-attempt.d.ts.map +1 -0
- package/dist/src/v2/active-attempt.js +200 -0
- package/dist/src/v2/active-attempt.js.map +1 -0
- package/dist/src/v2/adapters/command.d.ts +6 -0
- package/dist/src/v2/adapters/command.d.ts.map +1 -1
- package/dist/src/v2/adapters/command.js +43 -2
- package/dist/src/v2/adapters/command.js.map +1 -1
- package/dist/src/v2/candidate.d.ts +15 -31
- package/dist/src/v2/candidate.d.ts.map +1 -1
- package/dist/src/v2/candidate.js +7 -29
- package/dist/src/v2/candidate.js.map +1 -1
- package/dist/src/v2/checked-change.d.ts +3 -2
- package/dist/src/v2/checked-change.d.ts.map +1 -1
- package/dist/src/v2/checked-change.js +4 -3
- package/dist/src/v2/checked-change.js.map +1 -1
- package/dist/src/v2/cli-contract.d.ts +1 -1
- package/dist/src/v2/cli-contract.d.ts.map +1 -1
- package/dist/src/v2/cli-contract.js +4 -6
- package/dist/src/v2/cli-contract.js.map +1 -1
- package/dist/src/v2/cli.d.ts +8 -0
- package/dist/src/v2/cli.d.ts.map +1 -1
- package/dist/src/v2/cli.js +13 -0
- package/dist/src/v2/cli.js.map +1 -1
- package/dist/src/v2/code-review-report.d.ts +10 -18
- package/dist/src/v2/code-review-report.d.ts.map +1 -1
- package/dist/src/v2/code-review-report.js +63 -60
- package/dist/src/v2/code-review-report.js.map +1 -1
- package/dist/src/v2/codex-process.d.ts +6 -2
- package/dist/src/v2/codex-process.d.ts.map +1 -1
- package/dist/src/v2/codex-process.js +25 -9
- package/dist/src/v2/codex-process.js.map +1 -1
- package/dist/src/v2/config.d.ts +0 -2
- package/dist/src/v2/config.d.ts.map +1 -1
- package/dist/src/v2/config.js +3 -6
- package/dist/src/v2/config.js.map +1 -1
- package/dist/src/v2/contained-report-operation.d.ts +41 -196
- package/dist/src/v2/contained-report-operation.d.ts.map +1 -1
- package/dist/src/v2/contained-report-operation.js +139 -466
- package/dist/src/v2/contained-report-operation.js.map +1 -1
- package/dist/src/v2/containment.d.ts +1 -0
- package/dist/src/v2/containment.d.ts.map +1 -1
- package/dist/src/v2/containment.js +12 -2
- package/dist/src/v2/containment.js.map +1 -1
- package/dist/src/v2/delivery-authority.d.ts +26 -0
- package/dist/src/v2/delivery-authority.d.ts.map +1 -0
- package/dist/src/v2/delivery-authority.js +44 -0
- package/dist/src/v2/delivery-authority.js.map +1 -0
- package/dist/src/v2/direct-delivery.d.ts +16 -36
- package/dist/src/v2/direct-delivery.d.ts.map +1 -1
- package/dist/src/v2/direct-delivery.js +135 -122
- package/dist/src/v2/direct-delivery.js.map +1 -1
- package/dist/src/v2/immutable-workflow-publisher.d.ts.map +1 -1
- package/dist/src/v2/immutable-workflow-publisher.js +3 -1
- package/dist/src/v2/immutable-workflow-publisher.js.map +1 -1
- package/dist/src/v2/implementation-report.d.ts +3 -1
- package/dist/src/v2/implementation-report.d.ts.map +1 -1
- package/dist/src/v2/implementation-report.js +17 -4
- package/dist/src/v2/implementation-report.js.map +1 -1
- package/dist/src/v2/implementation-reviewer.d.ts +41 -12
- package/dist/src/v2/implementation-reviewer.d.ts.map +1 -1
- package/dist/src/v2/implementation-reviewer.js +114 -42
- package/dist/src/v2/implementation-reviewer.js.map +1 -1
- package/dist/src/v2/pending-effect-settlement.d.ts +44 -0
- package/dist/src/v2/pending-effect-settlement.d.ts.map +1 -0
- package/dist/src/v2/pending-effect-settlement.js +69 -0
- package/dist/src/v2/pending-effect-settlement.js.map +1 -0
- package/dist/src/v2/process-identity.d.ts +45 -0
- package/dist/src/v2/process-identity.d.ts.map +1 -0
- package/dist/src/v2/process-identity.js +118 -0
- package/dist/src/v2/process-identity.js.map +1 -0
- package/dist/src/v2/proof-report.d.ts +2 -1
- package/dist/src/v2/proof-report.d.ts.map +1 -1
- package/dist/src/v2/proof-report.js +10 -4
- package/dist/src/v2/proof-report.js.map +1 -1
- package/dist/src/v2/review-feedback-coordinator.d.ts +1 -1
- package/dist/src/v2/review-feedback-coordinator.d.ts.map +1 -1
- package/dist/src/v2/review-feedback-coordinator.js +1 -1
- package/dist/src/v2/review-feedback-coordinator.js.map +1 -1
- package/dist/src/v2/review-feedback.d.ts +14 -20
- package/dist/src/v2/review-feedback.d.ts.map +1 -1
- package/dist/src/v2/review-feedback.js +45 -87
- package/dist/src/v2/review-feedback.js.map +1 -1
- package/dist/src/v2/run-issue.d.ts +129 -88
- package/dist/src/v2/run-issue.d.ts.map +1 -1
- package/dist/src/v2/run-issue.js +1965 -2381
- package/dist/src/v2/run-issue.js.map +1 -1
- package/dist/src/v2/run-state-projections.d.ts +84 -0
- package/dist/src/v2/run-state-projections.d.ts.map +1 -0
- package/dist/src/v2/run-state-projections.js +142 -0
- package/dist/src/v2/run-state-projections.js.map +1 -0
- package/dist/src/v2/run-store.d.ts +99 -81
- package/dist/src/v2/run-store.d.ts.map +1 -1
- package/dist/src/v2/run-store.js +245 -542
- package/dist/src/v2/run-store.js.map +1 -1
- package/dist/src/v2/runtime-assets.d.ts +3 -0
- package/dist/src/v2/runtime-assets.d.ts.map +1 -1
- package/dist/src/v2/runtime-assets.js +104 -0
- package/dist/src/v2/runtime-assets.js.map +1 -1
- package/dist/src/v2/runtime.d.ts +56 -44
- package/dist/src/v2/runtime.d.ts.map +1 -1
- package/dist/src/v2/runtime.js +383 -503
- package/dist/src/v2/runtime.js.map +1 -1
- package/dist/src/v2/setup.js +0 -2
- package/dist/src/v2/setup.js.map +1 -1
- package/dist/src/v2/validation-progression.d.ts +70 -0
- package/dist/src/v2/validation-progression.d.ts.map +1 -0
- package/dist/src/v2/validation-progression.js +247 -0
- package/dist/src/v2/validation-progression.js.map +1 -0
- package/dist/src/v2/workflow-assets.d.ts +9 -3
- package/dist/src/v2/workflow-assets.d.ts.map +1 -1
- package/dist/src/v2/workflow-assets.js +256 -43
- package/dist/src/v2/workflow-assets.js.map +1 -1
- package/internal-workflow/docs/agents/bug-workflow-routing.md +9 -7
- package/internal-workflow/docs/agents/coding-skill-routing.md +170 -120
- package/internal-workflow/docs/agents/tool-usage.md +23 -12
- package/internal-workflow/manifest.json +1 -1
- package/internal-workflow/operations/code-review/SKILL.md +34 -15
- package/internal-workflow/operations/implementation/SKILL.md +21 -16
- package/internal-workflow/profiles/implementer.toml +9 -0
- package/internal-workflow/profiles/review_coordinator.toml +9 -0
- package/internal-workflow/profiles/spec_reviewer.toml +9 -0
- package/internal-workflow/profiles/standards_reviewer.toml +9 -0
- package/internal-workflow/schemas/code-review-v1.json +1 -1
- package/internal-workflow/schemas/implementation-report-v1.json +1 -1
- package/internal-workflow/schemas/proof-report-v1.json +1 -1
- package/internal-workflow/skills/bug-root-cause-explainer/SKILL.md +114 -0
- package/internal-workflow/skills/bug-root-cause-explainer/agents/openai.yaml +7 -0
- package/internal-workflow/skills/bug-root-cause-explainer/evals/evals.json +18 -0
- package/internal-workflow/skills/code-review/SKILL.md +84 -306
- package/internal-workflow/skills/code-review/agents/openai.yaml +5 -3
- package/internal-workflow/skills/code-review/evals/evals.json +83 -0
- package/internal-workflow/skills/code-review/references/standards-smells.md +41 -0
- package/internal-workflow/skills/diagnosing-bugs/SKILL.md +69 -32
- package/internal-workflow/skills/diagnosing-bugs/agents/openai.yaml +2 -2
- package/internal-workflow/skills/diagnosing-bugs/evals/evals.json +63 -0
- package/internal-workflow/skills/grilling/SKILL.md +51 -0
- package/internal-workflow/skills/grilling/agents/openai.yaml +6 -0
- package/internal-workflow/skills/grilling/evals/evals.json +47 -0
- package/internal-workflow/skills/implement/SKILL.md +135 -0
- package/internal-workflow/skills/implement/agents/openai.yaml +6 -0
- package/internal-workflow/skills/implement/evals/evals.json +150 -0
- package/internal-workflow/skills/plan/SKILL.md +59 -0
- package/internal-workflow/skills/plan/agents/openai.yaml +6 -0
- package/internal-workflow/skills/plan/evals/evals.json +36 -0
- package/internal-workflow/skills/prototype/LOGIC.md +130 -0
- package/internal-workflow/skills/prototype/SKILL.md +69 -0
- package/internal-workflow/skills/prototype/UI.md +157 -0
- package/internal-workflow/skills/prototype/agents/openai.yaml +6 -0
- package/internal-workflow/skills/prototype/evals/evals.json +67 -0
- package/internal-workflow/skills/research/SKILL.md +110 -0
- package/internal-workflow/skills/research/agents/openai.yaml +6 -0
- package/internal-workflow/skills/research/evals/evals.json +49 -0
- package/internal-workflow/skills/tdd/SKILL.md +72 -67
- package/internal-workflow/skills/tdd/agents/openai.yaml +2 -2
- package/internal-workflow/skills/tdd/evals/evals.json +12 -0
- package/internal-workflow/skills/tdd/mocking.md +48 -1
- package/internal-workflow/skills/tdd/refactoring.md +3 -3
- package/internal-workflow/skills/tickets-orchestrator/SKILL.md +199 -0
- package/internal-workflow/skills/tickets-orchestrator/agents/openai.yaml +6 -0
- package/internal-workflow/skills/tickets-orchestrator/evals/evals.json +126 -0
- package/internal-workflow/skills/tickets-orchestrator/references/delegate-integrate.md +83 -0
- package/internal-workflow/skills/tickets-orchestrator/references/finish-delivery.md +69 -0
- package/internal-workflow/skills/tickets-orchestrator/references/stop-completion.md +63 -0
- package/internal-workflow/skills/to-spec/SKILL.md +133 -0
- package/internal-workflow/skills/to-spec/agents/openai.yaml +6 -0
- package/internal-workflow/skills/to-spec/evals/evals.json +24 -0
- package/internal-workflow/skills/to-tickets/SKILL.md +189 -0
- package/internal-workflow/skills/to-tickets/agents/openai.yaml +6 -0
- package/internal-workflow/skills/to-tickets/evals/evals.json +79 -0
- package/internal-workflow/skills/to-tickets/references/publishing-details.md +117 -0
- package/package.json +1 -1
- package/dist/src/v2/proof-store.d.ts +0 -54
- package/dist/src/v2/proof-store.d.ts.map +0 -1
- package/dist/src/v2/proof-store.js +0 -301
- package/dist/src/v2/proof-store.js.map +0 -1
- package/dist/src/v2/route-continuations.d.ts +0 -32
- package/dist/src/v2/route-continuations.d.ts.map +0 -1
- package/dist/src/v2/route-continuations.js +0 -2
- package/dist/src/v2/route-continuations.js.map +0 -1
- package/dist/src/v2/route-coordinator.d.ts +0 -72
- package/dist/src/v2/route-coordinator.d.ts.map +0 -1
- package/dist/src/v2/route-coordinator.js +0 -275
- package/dist/src/v2/route-coordinator.js.map +0 -1
- package/dist/src/v2/route-decision.d.ts +0 -120
- package/dist/src/v2/route-decision.d.ts.map +0 -1
- package/dist/src/v2/route-decision.js +0 -380
- package/dist/src/v2/route-decision.js.map +0 -1
- package/dist/src/v2/spec-coordinator.d.ts +0 -73
- package/dist/src/v2/spec-coordinator.d.ts.map +0 -1
- package/dist/src/v2/spec-coordinator.js +0 -126
- package/dist/src/v2/spec-coordinator.js.map +0 -1
- package/dist/src/v2/spec-delivery.d.ts +0 -112
- package/dist/src/v2/spec-delivery.d.ts.map +0 -1
- package/dist/src/v2/spec-delivery.js +0 -336
- package/dist/src/v2/spec-delivery.js.map +0 -1
- package/dist/src/v2/triage-route.d.ts +0 -68
- package/dist/src/v2/triage-route.d.ts.map +0 -1
- package/dist/src/v2/triage-route.js +0 -223
- package/dist/src/v2/triage-route.js.map +0 -1
- package/dist/src/v2/waiting-human-coordinator.d.ts +0 -49
- package/dist/src/v2/waiting-human-coordinator.d.ts.map +0 -1
- package/dist/src/v2/waiting-human-coordinator.js +0 -509
- package/dist/src/v2/waiting-human-coordinator.js.map +0 -1
- package/dist/src/v2/waiting-human.d.ts +0 -143
- package/dist/src/v2/waiting-human.d.ts.map +0 -1
- package/dist/src/v2/waiting-human.js +0 -408
- package/dist/src/v2/waiting-human.js.map +0 -1
- package/internal-workflow/docs/agents/contract-test-ledger.md +0 -71
- package/internal-workflow/docs/agents/review-gates.md +0 -42
- package/internal-workflow/docs/agents/review-protocol.md +0 -98
- package/internal-workflow/evals/coding-skill-evals.json +0 -373
- package/internal-workflow/operations/ambiguity-review/SKILL.md +0 -5
- package/internal-workflow/operations/qualification-repair/SKILL.md +0 -17
- package/internal-workflow/operations/spec-author/SKILL.md +0 -12
- package/internal-workflow/operations/spec-review/SKILL.md +0 -12
- package/internal-workflow/operations/triage/SKILL.md +0 -12
- package/internal-workflow/profiles/analyst_deep.toml +0 -9
- package/internal-workflow/profiles/implementer_standard.toml +0 -9
- package/internal-workflow/profiles/proof_agent.toml +0 -8
- package/internal-workflow/profiles/reviewer_deep.toml +0 -9
- package/internal-workflow/profiles/reviewer_standard.toml +0 -9
- package/internal-workflow/schemas/ambiguity-review-v1.json +0 -1
- package/internal-workflow/schemas/spec-author-v1.json +0 -1
- package/internal-workflow/schemas/spec-review-v1.json +0 -30
- package/internal-workflow/schemas/triage-route-v1.json +0 -1
- package/internal-workflow/skills/agent-auto/SKILL.md +0 -19
- package/internal-workflow/skills/agent-auto/agents/openai.yaml +0 -6
- package/internal-workflow/skills/code-debugger/SKILL.md +0 -122
- package/internal-workflow/skills/code-debugger/agents/openai.yaml +0 -7
- package/internal-workflow/skills/code-review/references/bug-classes.md +0 -56
- package/internal-workflow/skills/code-review/references/cleanup-lens.md +0 -52
- package/internal-workflow/skills/code-review/references/framework-lenses.md +0 -34
- package/internal-workflow/skills/code-review/references/targeted-recipes.md +0 -49
- package/internal-workflow/skills/implementation-spec-maker/SKILL.md +0 -107
- package/internal-workflow/skills/implementation-spec-maker/agents/openai.yaml +0 -6
- package/internal-workflow/skills/implementation-spec-maker/references/source-modes.md +0 -32
- package/internal-workflow/skills/implementation-spec-maker/references/spec-template.md +0 -146
- package/internal-workflow/skills/implementation-spec-review/SKILL.md +0 -131
- package/internal-workflow/skills/implementation-spec-review/agents/openai.yaml +0 -6
- package/internal-workflow/skills/implementation-spec-review/evals/evals.json +0 -78
- package/internal-workflow/skills/implementation-spec-review/references/review-loop.md +0 -121
- package/internal-workflow/skills/small-task-implementer/SKILL.md +0 -112
- package/internal-workflow/skills/small-task-implementer/agents/openai.yaml +0 -6
- package/internal-workflow/skills/spec-implementer/SKILL.md +0 -133
- package/internal-workflow/skills/spec-implementer/agents/openai.yaml +0 -6
- package/internal-workflow/skills/spec-implementer/evals/evals.json +0 -30
- package/internal-workflow/skills/spec-implementer/references/review-loop.md +0 -100
- package/internal-workflow/skills/triage/AGENT-BRIEF.md +0 -192
- package/internal-workflow/skills/triage/OUT-OF-SCOPE.md +0 -101
- package/internal-workflow/skills/triage/SKILL.md +0 -134
- package/internal-workflow/skills/triage/agents/openai.yaml +0 -6
|
@@ -1,71 +0,0 @@
|
|
|
1
|
-
# Contract Test Ledger
|
|
2
|
-
|
|
3
|
-
Use this shared ledger for behavior-changing work where a passing happy-path test could still miss a contract defect. The ledger turns review-class risks into testable obligations before implementation.
|
|
4
|
-
|
|
5
|
-
## When Required
|
|
6
|
-
|
|
7
|
-
Create or update a ledger only when both are true:
|
|
8
|
-
|
|
9
|
-
1. The task materially changes an observable or shared contract.
|
|
10
|
-
2. You can name a realistic failure that ordinary targeted proof could miss.
|
|
11
|
-
|
|
12
|
-
Relevant contract areas include:
|
|
13
|
-
|
|
14
|
-
- API, DTO, schema, serialization, persistence, or externally visible response shape
|
|
15
|
-
- ordering, lifecycle events, state transitions, retries, idempotency, timeout, cancellation, or background jobs
|
|
16
|
-
- cache keys, invalidation, state merge precedence, fallback behavior, profile/global/mobile overrides, defaults, or feature flags
|
|
17
|
-
- evidence, trace, snapshot, audit, summary, aggregation, score, winner, or generated artifacts
|
|
18
|
-
- shared behavior read by multiple callers, tenants, users, groups, children, or projections
|
|
19
|
-
|
|
20
|
-
Category alone does not activate the ledger. Ordinary API/DTO edits use targeted
|
|
21
|
-
proof when no missed failure is named. Retry, idempotency, auth, payment,
|
|
22
|
-
shared-trust, and recovery changes usually pass the gate when happy-path proof
|
|
23
|
-
cannot expose their failure mode.
|
|
24
|
-
|
|
25
|
-
For narrow UI copy, docs-only, formatting, tests-only, or isolated styling changes, the ledger is not required.
|
|
26
|
-
|
|
27
|
-
## Required Shape
|
|
28
|
-
|
|
29
|
-
Keep the ledger compact. Use one row per invariant:
|
|
30
|
-
|
|
31
|
-
```markdown
|
|
32
|
-
## Contract Test Ledger
|
|
33
|
-
|
|
34
|
-
| Invariant | Risk It Prevents | First Test / Proof | Status |
|
|
35
|
-
| --- | --- | --- | --- |
|
|
36
|
-
| <observable rule> | <real failure mode> | <exact test name/command or manual proof> | planned / red / green / blocked |
|
|
37
|
-
```
|
|
38
|
-
|
|
39
|
-
Rules:
|
|
40
|
-
|
|
41
|
-
- The invariant must be observable through the public interface or the same seam real callers use.
|
|
42
|
-
- The risk must name the concrete bug class, not a vague "edge case".
|
|
43
|
-
- The first test/proof must fail before the fix unless the ledger records why a RED signal is impossible.
|
|
44
|
-
- `blocked` requires the missing seam, fixture, service, or decision that prevents proof.
|
|
45
|
-
- Keep the ledger current as implementation proceeds; do not backfill it only at the end.
|
|
46
|
-
|
|
47
|
-
## Invariant Prompts
|
|
48
|
-
|
|
49
|
-
Ask the relevant subset before the first RED test:
|
|
50
|
-
|
|
51
|
-
- **Ordering:** What must happen before/after terminal events, snapshots, persistence writes, notifications, or cleanup?
|
|
52
|
-
- **Precedence:** Which source wins among user input, profile, mobile, global, server, cache, AI, fallback, default, `null`, `false`, `0`, and empty objects?
|
|
53
|
-
- **Threading:** Does each new field survive construction, normalization, cloning, retry, persistence reload, serialization, and every visible consumer?
|
|
54
|
-
- **Runtime contract:** Do validation, internal types, persistence schema, API response, and consumers agree on names, units, nullability, enum values, and date/object formats?
|
|
55
|
-
- **Retry/idempotency:** What changes on retry, and which snapshots, counters, streams, timestamps, writes, or side effects must be rebuilt instead of reused?
|
|
56
|
-
- **Determinism:** When sort keys, timestamps, scores, priorities, or winners tie, what stable tie-breaker makes output repeatable?
|
|
57
|
-
- **Evidence:** Which trace, audit, snapshot, Fresh-Context, summary, or generated artifact proves the behavior actually happened?
|
|
58
|
-
- **Proof compatibility:** Can that source observe the exact claim at the required item/run/aggregate cardinality, and could it pass while the claim is false because of delay, redaction, correlation, or variant differences?
|
|
59
|
-
- **Partial failure:** If a dependency times out, throws, returns stale data, or fails after a side effect, what durable state remains and who repairs it?
|
|
60
|
-
- **Scope/cardinality:** Is data global, per-tenant, per-group, per-child, per-step, or per-item, and can one top-level field collapse multiple meaningful results?
|
|
61
|
-
|
|
62
|
-
## Review Feedback Loop
|
|
63
|
-
|
|
64
|
-
When code review finds a real contract defect, add or update one ledger row before fixing it:
|
|
65
|
-
|
|
66
|
-
- `Invariant`: the rule the implementation violated
|
|
67
|
-
- `Risk It Prevents`: the observed review finding
|
|
68
|
-
- `First Test / Proof`: the regression test or proof that would have caught it
|
|
69
|
-
- `Status`: `red` before the fix, then `green` after verification
|
|
70
|
-
|
|
71
|
-
If no correct public seam exists for the regression test, record that as `blocked` and name the architecture/testability gap. Do not replace a missing seam with an implementation-detail test unless the task explicitly approves that tradeoff.
|
|
@@ -1,42 +0,0 @@
|
|
|
1
|
-
# Review Gates
|
|
2
|
-
|
|
3
|
-
This file owns review applicability. Review execution mechanics live in
|
|
4
|
-
[`review-protocol.md`](review-protocol.md); approved-spec review shape lives in
|
|
5
|
-
[`spec-implementer/references/review-loop.md`](../../skills/spec-implementer/references/review-loop.md).
|
|
6
|
-
|
|
7
|
-
## Final Code Review
|
|
8
|
-
|
|
9
|
-
Run final `$code-review` for behavior-changing work that affects:
|
|
10
|
-
|
|
11
|
-
- medium/large shared business behavior;
|
|
12
|
-
- API/DTO/schema, migration, persistence, auth, permission, payment, cache,
|
|
13
|
-
concurrency, background jobs, or shared-state contracts;
|
|
14
|
-
- shared UI/navigation/middleware/core flows;
|
|
15
|
-
- runtime logic across three or more files when it crosses an owner or
|
|
16
|
-
validation seam.
|
|
17
|
-
|
|
18
|
-
Do not invoke it automatically for docs, copy, comments, tests-only or
|
|
19
|
-
styling-only changes, formatting, renames, mechanical refactors, or isolated
|
|
20
|
-
low-risk one-file fixes.
|
|
21
|
-
|
|
22
|
-
`medium` is the normal review profile. API, persistence, statefulness, file
|
|
23
|
-
count, or orchestration strengthens the review focus only when it creates an
|
|
24
|
-
affected contract; none independently selects `high`.
|
|
25
|
-
|
|
26
|
-
## Cleanup
|
|
27
|
-
|
|
28
|
-
Cleanup is a lens inside the same final `$code-review`, never a separate gate.
|
|
29
|
-
Use bounded cleanup by default. Amplify it only when the user, approved source,
|
|
30
|
-
or repository policy names a concrete evidenced simplification risk; follow
|
|
31
|
-
`../../skills/code-review/references/cleanup-lens.md` for that branch.
|
|
32
|
-
|
|
33
|
-
After one consolidated repair, coordinator verification plus affected
|
|
34
|
-
validation closes ordinary medium/low behavior-preserving findings. Use
|
|
35
|
-
Closure only for the triggers in `review-protocol.md`.
|
|
36
|
-
|
|
37
|
-
## Validation Depth
|
|
38
|
-
|
|
39
|
-
For simple and medium work, run targeted behavior proof and the smallest
|
|
40
|
-
affected integration check. Run a full repository suite only when explicitly
|
|
41
|
-
required by repository policy, when broad contract fan-out cannot be isolated,
|
|
42
|
-
or for a genuinely `high` task.
|
|
@@ -1,98 +0,0 @@
|
|
|
1
|
-
# Review Protocol
|
|
2
|
-
|
|
3
|
-
This file owns mechanics shared by artifact and implementation review: bounded
|
|
4
|
-
capsules, Full/Closure, defect lifecycle, no-progress, waiver, and the result
|
|
5
|
-
envelope. Target Modules own authority, profile, topology, durable state, and
|
|
6
|
-
outcome mapping.
|
|
7
|
-
|
|
8
|
-
## Capsule
|
|
9
|
-
|
|
10
|
-
Every reviewer receives only:
|
|
11
|
-
|
|
12
|
-
- target kind/path and pinned revision;
|
|
13
|
-
- one unanswered review question and assigned lenses;
|
|
14
|
-
- authority, approved scope, and relevant evidence;
|
|
15
|
-
- compact affected validation and current defect records;
|
|
16
|
-
- for Closure, the repair diff and mapping from changed targets to defects and
|
|
17
|
-
affected contracts.
|
|
18
|
-
|
|
19
|
-
Do not pass raw parent history, old revisions, repeated logs, or unrelated
|
|
20
|
-
inventories. Reuse valid coverage for the same revision, question, and lenses.
|
|
21
|
-
|
|
22
|
-
## Modes
|
|
23
|
-
|
|
24
|
-
**Full** covers the complete assigned scope once. It is bounded to the settled
|
|
25
|
-
target, changed owners, authority, and plausible affected contracts. It is not
|
|
26
|
-
a repository-wide audit and does not activate unrelated lenses.
|
|
27
|
-
|
|
28
|
-
**Closure** verifies a repair only when the finding is critical/high, affects a
|
|
29
|
-
trust boundary, durable data, concurrency/idempotency, shared API/DTO/schema, or
|
|
30
|
-
invalidates mandatory coverage. Closure stays with affected reviewer lineages
|
|
31
|
-
and targets. Do not launch it merely because Full found an ordinary defect.
|
|
32
|
-
|
|
33
|
-
A repair starts another Full only when mandatory-lens coverage became invalid.
|
|
34
|
-
A live reviewer poll timeout is non-terminal and does not authorize duplicate
|
|
35
|
-
review or cancellation; reconcile the recorded session first.
|
|
36
|
-
|
|
37
|
-
## Defects
|
|
38
|
-
|
|
39
|
-
Use one canonical record per distinct invariant and failure mechanism:
|
|
40
|
-
|
|
41
|
-
```yaml
|
|
42
|
-
id: REVIEW-CONC-003
|
|
43
|
-
class: blocker | execution-risk | improvement
|
|
44
|
-
status: open | fixed | verified | blocked | accepted-risk | superseded
|
|
45
|
-
severity: critical | high | medium | low
|
|
46
|
-
confidence: high | medium | low
|
|
47
|
-
invariant: "<observable rule>"
|
|
48
|
-
failure: "<concrete trigger and impact>"
|
|
49
|
-
evidence: ["<target or source>"]
|
|
50
|
-
repair: "<smallest sufficient change>"
|
|
51
|
-
affected_targets: ["<path, section, contract, or lens>"]
|
|
52
|
-
```
|
|
53
|
-
|
|
54
|
-
Root assigns stable IDs and deduplicates only when both invariant and failure
|
|
55
|
-
match. A superseded record must point to a distinct canonical replacement.
|
|
56
|
-
Improvements never block. Only an execution risk may become `accepted-risk`,
|
|
57
|
-
and only with explicit authority, reason, scope, and target revision. A blocker
|
|
58
|
-
cannot be accepted or downgraded.
|
|
59
|
-
|
|
60
|
-
Move blocking defects `open -> fixed -> verified`. For ordinary medium/low
|
|
61
|
-
behavior-preserving repairs, root may verify after checking the cited failure
|
|
62
|
-
path and affected validation. Closure-triggering defects require affected
|
|
63
|
-
independent verification.
|
|
64
|
-
|
|
65
|
-
In Closure, copy each supplied canonical defect's `id`, `class`, `invariant`,
|
|
66
|
-
`failure`, and introduced target revision byte-for-byte. Never paraphrase those
|
|
67
|
-
immutable fields while describing verification. Record the Closure decision in
|
|
68
|
-
the status, status target revision, evidence, and repair-finding outcome fields.
|
|
69
|
-
|
|
70
|
-
## Repair And Stop
|
|
71
|
-
|
|
72
|
-
Repair compatible findings in one consolidated batch. Before repeating review,
|
|
73
|
-
the target, evidence, repair, or source decision must change materially.
|
|
74
|
-
Review count and elapsed time are audit signals, never approval or blocking
|
|
75
|
-
conditions.
|
|
76
|
-
|
|
77
|
-
Stop when repair requires a product/scope/owner decision, mandatory evidence or
|
|
78
|
-
reviewer is unavailable, no substantive repair exists, or the same failure
|
|
79
|
-
repeats without progress. Do not create micro-cycles for ordinary medium/low
|
|
80
|
-
findings.
|
|
81
|
-
|
|
82
|
-
Waive review only after explicit user instruction. Record skipped coverage and
|
|
83
|
-
open defects. Waiver is not approval and accepts no defect automatically.
|
|
84
|
-
|
|
85
|
-
## Result
|
|
86
|
-
|
|
87
|
-
```text
|
|
88
|
-
Review Mode: <Full | Closure>
|
|
89
|
-
Mandatory Coverage: <covered lenses or gaps>
|
|
90
|
-
Verified Defects: <IDs or None>
|
|
91
|
-
Accepted Risks: <IDs, authority, reason or None>
|
|
92
|
-
Open Defects: <IDs or None>
|
|
93
|
-
```
|
|
94
|
-
|
|
95
|
-
Target Modules add profile, authority, outcome, and checkpoint without
|
|
96
|
-
redefining these fields. For a normal medium flow, do not report internal
|
|
97
|
-
session accounting unless interruption, Closure, accepted risk, or another
|
|
98
|
-
exception makes it relevant.
|
|
@@ -1,373 +0,0 @@
|
|
|
1
|
-
{
|
|
2
|
-
"schema_version": 1,
|
|
3
|
-
"purpose": "Cross-skill routing and handoff regressions only; skill-local behavior belongs in each owner's evals directory.",
|
|
4
|
-
"cases": [
|
|
5
|
-
{
|
|
6
|
-
"id": "route-feature-tdd",
|
|
7
|
-
"prompt": "Implement a behavior-changing feature in this repository.",
|
|
8
|
-
"expected": ["read local evidence", "activate tdd before behavior edits", "run affected validation"],
|
|
9
|
-
"forbidden": ["invent repository commands", "create planning artifacts without an execution gap"],
|
|
10
|
-
"execution": {
|
|
11
|
-
"level": "trigger",
|
|
12
|
-
"fixture": "routing/behavior-change",
|
|
13
|
-
"sandbox": "read-only",
|
|
14
|
-
"trials": 2,
|
|
15
|
-
"should_trigger": ["code-review", "tdd"],
|
|
16
|
-
"should_not_trigger": ["implementation-spec-maker", "small-task-implementer", "tickets-orchestrator"],
|
|
17
|
-
"assertions": [{"kind": "route_equals", "value": "direct-medium"}],
|
|
18
|
-
"ablation_skill": "tdd"
|
|
19
|
-
}
|
|
20
|
-
},
|
|
21
|
-
{
|
|
22
|
-
"id": "route-diagnosis-only",
|
|
23
|
-
"prompt": "Explain the confirmed root cause documented in this repository and do not fix it yet.",
|
|
24
|
-
"expected": ["use bug-root-cause-explainer", "cite concrete evidence", "stop before edits"],
|
|
25
|
-
"forbidden": ["edit files", "present an inference as proven"],
|
|
26
|
-
"execution": {
|
|
27
|
-
"level": "behavior",
|
|
28
|
-
"fixture": "behavior/diagnosis-only",
|
|
29
|
-
"sandbox": "workspace-write",
|
|
30
|
-
"trials": 2,
|
|
31
|
-
"should_trigger": ["bug-root-cause-explainer"],
|
|
32
|
-
"should_not_trigger": ["code-debugger", "diagnosing-bugs"],
|
|
33
|
-
"assertions": [
|
|
34
|
-
{"kind": "route_equals", "value": "diagnosis-only"},
|
|
35
|
-
{"kind": "tree_unchanged"}
|
|
36
|
-
],
|
|
37
|
-
"ablation_skill": "bug-root-cause-explainer"
|
|
38
|
-
}
|
|
39
|
-
},
|
|
40
|
-
{
|
|
41
|
-
"id": "route-flaky-bug",
|
|
42
|
-
"prompt": "Debug the intermittent failure described by this repository; its cause is not confirmed.",
|
|
43
|
-
"expected": ["use diagnosing-bugs", "improve reproduction or signal before repair"],
|
|
44
|
-
"forbidden": ["jump directly to a fix"],
|
|
45
|
-
"execution": {
|
|
46
|
-
"level": "trigger",
|
|
47
|
-
"fixture": "routing/flaky-bug",
|
|
48
|
-
"sandbox": "read-only",
|
|
49
|
-
"trials": 2,
|
|
50
|
-
"should_trigger": ["diagnosing-bugs"],
|
|
51
|
-
"should_not_trigger": ["bug-root-cause-explainer", "code-debugger"],
|
|
52
|
-
"assertions": [{"kind": "route_equals", "value": "diagnose-first"}],
|
|
53
|
-
"ablation_skill": "diagnosing-bugs"
|
|
54
|
-
}
|
|
55
|
-
},
|
|
56
|
-
{
|
|
57
|
-
"id": "profile-medium-default",
|
|
58
|
-
"prompt": "Implement the settled stateful API and persistence feature described in this repository.",
|
|
59
|
-
"expected": ["select medium", "implement directly when authority and proof are clear"],
|
|
60
|
-
"forbidden": ["select high from statefulness or file count", "manufacture a PRD, tickets, or spec"],
|
|
61
|
-
"execution": {
|
|
62
|
-
"level": "trigger",
|
|
63
|
-
"fixture": "routing/direct-medium",
|
|
64
|
-
"sandbox": "read-only",
|
|
65
|
-
"trials": 2,
|
|
66
|
-
"should_trigger": ["tdd"],
|
|
67
|
-
"should_not_trigger": ["implementation-spec-maker", "tickets-orchestrator"],
|
|
68
|
-
"assertions": [{"kind": "route_equals", "value": "direct-medium"}]
|
|
69
|
-
}
|
|
70
|
-
},
|
|
71
|
-
{
|
|
72
|
-
"id": "profile-high-threshold",
|
|
73
|
-
"prompt": "Classify the ordinary refactor and the irreversible data-loss scenario documented in this repository.",
|
|
74
|
-
"expected": ["keep the broad ordinary refactor medium", "select high only for material consequence plus uncertainty"],
|
|
75
|
-
"forbidden": ["accumulate generic risk labels into high"],
|
|
76
|
-
"execution": {
|
|
77
|
-
"level": "trigger",
|
|
78
|
-
"fixture": "routing/profile-classification",
|
|
79
|
-
"sandbox": "read-only",
|
|
80
|
-
"trials": 2,
|
|
81
|
-
"should_trigger": [],
|
|
82
|
-
"should_not_trigger": ["implementation-spec-maker", "tickets-orchestrator"],
|
|
83
|
-
"assertions": [{"kind": "route_equals", "value": "profile-classification"}]
|
|
84
|
-
}
|
|
85
|
-
},
|
|
86
|
-
{
|
|
87
|
-
"id": "planning-does-not-deliver",
|
|
88
|
-
"prompt": "Use wayfinder to resolve the documented decision frontier and publish the resulting tickets, then stop before delivery.",
|
|
89
|
-
"expected": ["stop after the planning package", "require separate delivery authorization"],
|
|
90
|
-
"forbidden": ["start implementation from publication or labels"],
|
|
91
|
-
"execution": {
|
|
92
|
-
"level": "trigger",
|
|
93
|
-
"fixture": "routing/planning",
|
|
94
|
-
"sandbox": "read-only",
|
|
95
|
-
"trials": 2,
|
|
96
|
-
"should_trigger": ["to-tickets", "wayfinder"],
|
|
97
|
-
"should_not_trigger": ["spec-implementer", "tickets-orchestrator"],
|
|
98
|
-
"assertions": [{"kind": "route_equals", "value": "planning-only"}],
|
|
99
|
-
"ablation_skill": "wayfinder"
|
|
100
|
-
}
|
|
101
|
-
},
|
|
102
|
-
{
|
|
103
|
-
"id": "spec-only-for-gap",
|
|
104
|
-
"prompt": "Choose the route for the implementation request whose unresolved execution contract is documented in this repository.",
|
|
105
|
-
"expected": ["use implementation-spec-maker for the unresolved execution delta"],
|
|
106
|
-
"forbidden": ["start implementation before the contract is resolved", "create a ticket graph"],
|
|
107
|
-
"execution": {
|
|
108
|
-
"level": "trigger",
|
|
109
|
-
"fixture": "routing/spec-gap",
|
|
110
|
-
"sandbox": "read-only",
|
|
111
|
-
"trials": 2,
|
|
112
|
-
"should_trigger": ["implementation-spec-maker"],
|
|
113
|
-
"should_not_trigger": ["spec-implementer", "tickets-orchestrator"],
|
|
114
|
-
"assertions": [{"kind": "route_equals", "value": "spec-gap"}],
|
|
115
|
-
"ablation_skill": "implementation-spec-maker"
|
|
116
|
-
}
|
|
117
|
-
},
|
|
118
|
-
{
|
|
119
|
-
"id": "approved-spec-execution",
|
|
120
|
-
"prompt": "Execute the approved implementation spec in this repository and stop at its stated delivery boundary.",
|
|
121
|
-
"expected": ["use spec-implementer", "follow the approved checklist through observable proof", "update reached checklist items truthfully"],
|
|
122
|
-
"forbidden": ["create a replacement spec", "route through tickets", "commit or push"],
|
|
123
|
-
"execution": {
|
|
124
|
-
"level": "end_to_end",
|
|
125
|
-
"fixture": "end-to-end-approved-spec",
|
|
126
|
-
"sandbox": "workspace-write",
|
|
127
|
-
"trials": 2,
|
|
128
|
-
"should_trigger": ["spec-implementer", "tdd"],
|
|
129
|
-
"should_not_trigger": ["implementation-spec-maker", "tickets-orchestrator"],
|
|
130
|
-
"assertions": [
|
|
131
|
-
{"kind": "route_equals", "value": "spec-implementer"},
|
|
132
|
-
{"kind": "file_contains", "path": ".eval/red.json", "text": "\"status\": \"red\""},
|
|
133
|
-
{"kind": "command_exit", "argv": ["python3", "verify_outcome.py"], "exit_code": 0, "boundary": false},
|
|
134
|
-
{"kind": "file_contains", "path": "docs/implementation-specs/approved-greeting.md", "text": "- [x] Implement the greeting"},
|
|
135
|
-
{"kind": "file_contains", "path": "docs/implementation-specs/approved-greeting.md", "text": "- [x] Run the focused test"},
|
|
136
|
-
{"kind": "event_present", "event": "file_write", "value": "greeting.py"},
|
|
137
|
-
{"kind": "event_order", "before": {"event": "skill_read", "value": "spec-implementer"}, "after": {"event": "file_write", "value": "greeting.py"}},
|
|
138
|
-
{"kind": "event_absent", "event": "git_push", "value": "origin"}
|
|
139
|
-
],
|
|
140
|
-
"ablation_skill": "spec-implementer"
|
|
141
|
-
}
|
|
142
|
-
},
|
|
143
|
-
{
|
|
144
|
-
"id": "ticket-direct-vs-graph",
|
|
145
|
-
"prompt": "Deliver the approved dependency graph documented in this repository through its public tests and stop at the local delivery boundary.",
|
|
146
|
-
"expected": ["use tickets-orchestrator for the graph", "deliver independent tickets before their dependent", "prove the completed graph"],
|
|
147
|
-
"forbidden": ["manufacture a wave-level spec", "skip a dependency", "commit or push"],
|
|
148
|
-
"execution": {
|
|
149
|
-
"level": "end_to_end",
|
|
150
|
-
"fixture": "end-to-end-ticket-graph",
|
|
151
|
-
"sandbox": "workspace-write",
|
|
152
|
-
"trials": 2,
|
|
153
|
-
"should_trigger": ["tdd", "tickets-orchestrator"],
|
|
154
|
-
"should_not_trigger": ["implementation-spec-maker", "spec-implementer"],
|
|
155
|
-
"assertions": [
|
|
156
|
-
{"kind": "route_equals", "value": "ticket-graph"},
|
|
157
|
-
{"kind": "file_contains", "path": ".eval/red.json", "text": "\"status\": \"red\""},
|
|
158
|
-
{"kind": "command_exit", "argv": ["python3", "verify_outcome.py"], "exit_code": 0, "boundary": false},
|
|
159
|
-
{"kind": "file_contains", "path": "delivery-state.md", "text": "- [x] Ticket A"},
|
|
160
|
-
{"kind": "file_contains", "path": "delivery-state.md", "text": "- [x] Ticket B"},
|
|
161
|
-
{"kind": "file_contains", "path": "delivery-state.md", "text": "- [x] Ticket C"},
|
|
162
|
-
{"kind": "event_present", "event": "file_write", "value": "summary.py"},
|
|
163
|
-
{"kind": "event_order", "before": {"event": "skill_read", "value": "tickets-orchestrator"}, "after": {"event": "file_write", "value": "summary.py"}},
|
|
164
|
-
{"kind": "path_absent", "path": "PLAN.md"},
|
|
165
|
-
{"kind": "event_absent", "event": "git_push", "value": "origin"}
|
|
166
|
-
],
|
|
167
|
-
"ablation_skill": "tickets-orchestrator"
|
|
168
|
-
}
|
|
169
|
-
},
|
|
170
|
-
{
|
|
171
|
-
"id": "review-topology",
|
|
172
|
-
"prompt": "Run the final review for the ordinary medium change documented in this repository.",
|
|
173
|
-
"expected": ["launch one reviewer_standard", "keep correctness and standards in one bounded Full review"],
|
|
174
|
-
"forbidden": ["root self-review", "launch reviewer_deep for ordinary medium"],
|
|
175
|
-
"execution": {
|
|
176
|
-
"level": "behavior",
|
|
177
|
-
"fixture": "behavior/review",
|
|
178
|
-
"sandbox": "workspace-write",
|
|
179
|
-
"trials": 2,
|
|
180
|
-
"should_trigger": ["code-review"],
|
|
181
|
-
"should_not_trigger": ["implementation-spec-review"],
|
|
182
|
-
"assertions": [
|
|
183
|
-
{"kind": "route_equals", "value": "review-topology"},
|
|
184
|
-
{"kind": "event_present", "event": "subagent_launch", "value": "reviewer_standard"},
|
|
185
|
-
{"kind": "event_absent", "event": "subagent_launch", "value": "reviewer_deep"},
|
|
186
|
-
{"kind": "tree_unchanged"}
|
|
187
|
-
]
|
|
188
|
-
}
|
|
189
|
-
},
|
|
190
|
-
{
|
|
191
|
-
"id": "review-closure-bounded",
|
|
192
|
-
"prompt": "Execute the already-authorized Closure for the repaired shared-contract defect documented in this repository.",
|
|
193
|
-
"expected": ["launch the recorded reviewer_standard lineage for affected Closure", "leave the already-repaired tree unchanged"],
|
|
194
|
-
"forbidden": ["restart broad Full review", "launch reviewer_deep or a separate cleanup gate"],
|
|
195
|
-
"execution": {
|
|
196
|
-
"level": "behavior",
|
|
197
|
-
"fixture": "behavior/review-closure",
|
|
198
|
-
"sandbox": "workspace-write",
|
|
199
|
-
"trials": 2,
|
|
200
|
-
"should_trigger": ["code-review"],
|
|
201
|
-
"should_not_trigger": ["implementation-spec-review"],
|
|
202
|
-
"assertions": [
|
|
203
|
-
{"kind": "route_equals", "value": "review-closure"},
|
|
204
|
-
{"kind": "event_present", "event": "subagent_launch", "value": "reviewer_standard:closure"},
|
|
205
|
-
{"kind": "event_absent", "event": "subagent_launch", "value": "reviewer_deep"},
|
|
206
|
-
{"kind": "tree_unchanged"}
|
|
207
|
-
]
|
|
208
|
-
}
|
|
209
|
-
},
|
|
210
|
-
{
|
|
211
|
-
"id": "ledger-needs-material-missed-failure",
|
|
212
|
-
"prompt": "Choose proof for the ordinary DTO edit documented in this repository; its targeted caller test covers the only realistic failure.",
|
|
213
|
-
"expected": ["use targeted proof without a contract ledger"],
|
|
214
|
-
"forbidden": ["create a ledger from the DTO category alone"],
|
|
215
|
-
"execution": {
|
|
216
|
-
"level": "trigger",
|
|
217
|
-
"fixture": "routing/ledger",
|
|
218
|
-
"sandbox": "read-only",
|
|
219
|
-
"trials": 2,
|
|
220
|
-
"should_trigger": ["tdd"],
|
|
221
|
-
"should_not_trigger": ["implementation-spec-maker"],
|
|
222
|
-
"assertions": [{"kind": "route_equals", "value": "targeted-proof"}]
|
|
223
|
-
}
|
|
224
|
-
},
|
|
225
|
-
{
|
|
226
|
-
"id": "ledger-protects-retry-contract",
|
|
227
|
-
"prompt": "Change the payment retry behavior documented in this repository, where a happy-path test could miss duplicate side effects.",
|
|
228
|
-
"expected": ["create a contract ledger", "name duplicate payment as the missed failure"],
|
|
229
|
-
"forbidden": ["rely on the happy path alone"],
|
|
230
|
-
"execution": {
|
|
231
|
-
"level": "trigger",
|
|
232
|
-
"fixture": "routing/ledger",
|
|
233
|
-
"sandbox": "read-only",
|
|
234
|
-
"trials": 2,
|
|
235
|
-
"should_trigger": ["tdd"],
|
|
236
|
-
"should_not_trigger": ["implementation-spec-maker"],
|
|
237
|
-
"assertions": [{"kind": "route_equals", "value": "contract-ledger"}]
|
|
238
|
-
}
|
|
239
|
-
},
|
|
240
|
-
{
|
|
241
|
-
"id": "direct-medium-no-extra-route",
|
|
242
|
-
"prompt": "Implement the clear medium feature documented in this repository; behavior, ownership, and affected proof are settled.",
|
|
243
|
-
"expected": ["implement directly"],
|
|
244
|
-
"forbidden": ["create tickets, a spec, or orchestration from file count, statefulness, or generic risk"],
|
|
245
|
-
"execution": {
|
|
246
|
-
"level": "end_to_end",
|
|
247
|
-
"fixture": "end-to-end-direct-medium",
|
|
248
|
-
"sandbox": "workspace-write",
|
|
249
|
-
"trials": 2,
|
|
250
|
-
"should_trigger": ["code-review", "tdd"],
|
|
251
|
-
"should_not_trigger": ["implementation-spec-maker", "small-task-implementer", "tickets-orchestrator"],
|
|
252
|
-
"assertions": [
|
|
253
|
-
{"kind": "route_equals", "value": "direct-medium"},
|
|
254
|
-
{"kind": "file_contains", "path": ".eval/red.json", "text": "\"status\": \"red\""},
|
|
255
|
-
{"kind": "command_exit", "argv": ["python3", "verify_outcome.py"], "exit_code": 0, "boundary": false},
|
|
256
|
-
{"kind": "path_absent", "path": "PLAN.md"},
|
|
257
|
-
{"kind": "path_absent", "path": "docs/implementation-specs"},
|
|
258
|
-
{"kind": "event_present", "event": "file_write", "value": "store.py"},
|
|
259
|
-
{"kind": "event_order", "before": {"event": "skill_read", "value": "tdd"}, "after": {"event": "file_write", "value": "store.py"}},
|
|
260
|
-
{"kind": "event_present", "event": "subagent_launch", "value": "reviewer_standard"},
|
|
261
|
-
{"kind": "event_absent", "event": "git_push", "value": "origin"}
|
|
262
|
-
]
|
|
263
|
-
}
|
|
264
|
-
},
|
|
265
|
-
{
|
|
266
|
-
"id": "route-copy-config",
|
|
267
|
-
"prompt": "Apply the exact copy and configuration-only update documented in this repository.",
|
|
268
|
-
"expected": ["use the small-task route", "run affected validation without manufacturing RED"],
|
|
269
|
-
"forbidden": ["activate tdd", "create an implementation spec"],
|
|
270
|
-
"execution": {
|
|
271
|
-
"level": "trigger",
|
|
272
|
-
"fixture": "routing/nonbehavior",
|
|
273
|
-
"sandbox": "read-only",
|
|
274
|
-
"trials": 2,
|
|
275
|
-
"should_trigger": ["small-task-implementer"],
|
|
276
|
-
"should_not_trigger": ["implementation-spec-maker", "tdd"],
|
|
277
|
-
"assertions": [{"kind": "route_equals", "value": "direct-simple"}],
|
|
278
|
-
"ablation_skill": "small-task-implementer"
|
|
279
|
-
}
|
|
280
|
-
},
|
|
281
|
-
{
|
|
282
|
-
"id": "route-bounded-fix",
|
|
283
|
-
"prompt": "Implement the confirmed bounded bug fix documented in this repository through its existing public test seam.",
|
|
284
|
-
"expected": ["activate tdd", "use code-debugger", "run affected validation"],
|
|
285
|
-
"forbidden": ["restart open-ended diagnosis", "create an implementation spec"],
|
|
286
|
-
"execution": {
|
|
287
|
-
"level": "end_to_end",
|
|
288
|
-
"fixture": "end-to-end-bounded-fix",
|
|
289
|
-
"sandbox": "workspace-write",
|
|
290
|
-
"trials": 2,
|
|
291
|
-
"should_trigger": ["code-debugger", "tdd"],
|
|
292
|
-
"should_not_trigger": ["diagnosing-bugs", "implementation-spec-maker"],
|
|
293
|
-
"assertions": [
|
|
294
|
-
{"kind": "route_equals", "value": "bounded-fix"},
|
|
295
|
-
{"kind": "file_contains", "path": ".eval/red.json", "text": "\"status\": \"red\""},
|
|
296
|
-
{"kind": "command_exit", "argv": ["python3", "verify_outcome.py"], "exit_code": 0, "boundary": false},
|
|
297
|
-
{"kind": "event_present", "event": "file_write", "value": "calculator.py"},
|
|
298
|
-
{"kind": "event_order", "before": {"event": "skill_read", "value": "tdd"}, "after": {"event": "file_write", "value": "calculator.py"}},
|
|
299
|
-
{"kind": "path_absent", "path": "PLAN.md"},
|
|
300
|
-
{"kind": "event_absent", "event": "git_push", "value": "origin"}
|
|
301
|
-
],
|
|
302
|
-
"ablation_skill": "code-debugger"
|
|
303
|
-
}
|
|
304
|
-
},
|
|
305
|
-
{
|
|
306
|
-
"id": "route-commit-only",
|
|
307
|
-
"prompt": "Commit only the already validated scoped changes described in this repository and stop without pushing.",
|
|
308
|
-
"expected": ["use commit", "stage only intended paths", "stop after the local commit"],
|
|
309
|
-
"forbidden": ["push", "open a pull request", "start another implementation workflow"],
|
|
310
|
-
"execution": {
|
|
311
|
-
"level": "behavior",
|
|
312
|
-
"fixture": "behavior/commit-only",
|
|
313
|
-
"sandbox": "workspace-write",
|
|
314
|
-
"trials": 2,
|
|
315
|
-
"should_trigger": ["commit"],
|
|
316
|
-
"should_not_trigger": ["code-review", "tickets-orchestrator"],
|
|
317
|
-
"assertions": [
|
|
318
|
-
{"kind": "route_equals", "value": "commit-only"},
|
|
319
|
-
{"kind": "command_exit", "argv": ["python3", "verify_commit.py"], "exit_code": 0, "boundary": true},
|
|
320
|
-
{"kind": "event_present", "event": "git_commit", "value": "commit"},
|
|
321
|
-
{"kind": "event_absent", "event": "git_push", "value": "origin"},
|
|
322
|
-
{"kind": "tree_unchanged"}
|
|
323
|
-
],
|
|
324
|
-
"ablation_skill": "commit"
|
|
325
|
-
}
|
|
326
|
-
},
|
|
327
|
-
{
|
|
328
|
-
"id": "route-evidence-answers",
|
|
329
|
-
"prompt": "Implement the tiny change using the exact target, expected value, and validation command already documented in this repository.",
|
|
330
|
-
"expected": ["use repository evidence", "continue without an obvious clarification question"],
|
|
331
|
-
"forbidden": ["ask for details already present", "create planning artifacts"],
|
|
332
|
-
"execution": {
|
|
333
|
-
"level": "trigger",
|
|
334
|
-
"fixture": "routing/evidence-first",
|
|
335
|
-
"sandbox": "read-only",
|
|
336
|
-
"trials": 2,
|
|
337
|
-
"should_trigger": ["small-task-implementer"],
|
|
338
|
-
"should_not_trigger": ["implementation-spec-maker", "wayfinder"],
|
|
339
|
-
"assertions": [{"kind": "route_equals", "value": "evidence-first"}]
|
|
340
|
-
}
|
|
341
|
-
},
|
|
342
|
-
{
|
|
343
|
-
"id": "route-invented-defaults",
|
|
344
|
-
"prompt": "Prepare the implementation route for the configurable lookback and cooldown requirement documented in this repository.",
|
|
345
|
-
"expected": ["identify the missing authority values", "use implementation-spec-maker and block the unresolved decision"],
|
|
346
|
-
"forbidden": ["invent defaults", "start implementation"],
|
|
347
|
-
"execution": {
|
|
348
|
-
"level": "trigger",
|
|
349
|
-
"fixture": "routing/authority-gap",
|
|
350
|
-
"sandbox": "read-only",
|
|
351
|
-
"trials": 2,
|
|
352
|
-
"should_trigger": ["implementation-spec-maker"],
|
|
353
|
-
"should_not_trigger": ["small-task-implementer", "spec-implementer"],
|
|
354
|
-
"assertions": [{"kind": "route_equals", "value": "authority-blocked"}]
|
|
355
|
-
}
|
|
356
|
-
},
|
|
357
|
-
{
|
|
358
|
-
"id": "route-preserve-existing",
|
|
359
|
-
"prompt": "Prepare an implementation spec for the request documented in this repository, preserving behavior that repository evidence proves already exists.",
|
|
360
|
-
"expected": ["classify existing behavior as preserve plus regression proof", "scope only the remaining delta"],
|
|
361
|
-
"forbidden": ["plan a second implementation", "ignore current code evidence"],
|
|
362
|
-
"execution": {
|
|
363
|
-
"level": "trigger",
|
|
364
|
-
"fixture": "routing/preserve-existing",
|
|
365
|
-
"sandbox": "read-only",
|
|
366
|
-
"trials": 2,
|
|
367
|
-
"should_trigger": ["implementation-spec-maker"],
|
|
368
|
-
"should_not_trigger": ["code-debugger", "tdd"],
|
|
369
|
-
"assertions": [{"kind": "route_equals", "value": "preserve-existing"}]
|
|
370
|
-
}
|
|
371
|
-
}
|
|
372
|
-
]
|
|
373
|
-
}
|
|
@@ -1,5 +0,0 @@
|
|
|
1
|
-
# Fresh Ambiguity Review Operation
|
|
2
|
-
|
|
3
|
-
Independently verify that a supplied waiting candidate contains at least two materially different product outcomes and no source-authorized choice. Technical, architecture, test, and tool choices are never product ambiguity. Do not edit files or external state. Return only `schemas/ambiguity-review-v1.json`.
|
|
4
|
-
|
|
5
|
-
Apply the packaged [confidence rubric](../../docs/agents/confidence-rubric.md).
|
|
@@ -1,17 +0,0 @@
|
|
|
1
|
-
# Qualification Repair Operation
|
|
2
|
-
|
|
3
|
-
Repair only the scoped check failures supplied by the Runner. This operation
|
|
4
|
-
does not authorize implementation of the issue acceptance criteria. Use the
|
|
5
|
-
smallest relevant debugging or TDD workflow needed to make the supplied checks
|
|
6
|
-
pass, and do not broaden validation beyond those commands.
|
|
7
|
-
|
|
8
|
-
Use [Code Debugger](../../skills/code-debugger/SKILL.md) for a confirmed check
|
|
9
|
-
failure, [Diagnosing Bugs](../../skills/diagnosing-bugs/SKILL.md) only when its
|
|
10
|
-
cause is unclear, and [TDD](../../skills/tdd/SKILL.md) only when the repair
|
|
11
|
-
changes observable behavior.
|
|
12
|
-
|
|
13
|
-
The Runner owns issue implementation, review, final checks, commits,
|
|
14
|
-
publication, retries, and external state. Never commit, push, publish, mutate
|
|
15
|
-
GitHub, or expose credentials. In the final report, `changedFiles` is the
|
|
16
|
-
complete current worktree change set, including changes that existed before
|
|
17
|
-
this attempt. Return only `schemas/implementation-report-v1.json`.
|
|
@@ -1,12 +0,0 @@
|
|
|
1
|
-
# Spec Author Operation
|
|
2
|
-
|
|
3
|
-
Follow the authoring and minimum-solution contract in the packaged
|
|
4
|
-
[Implementation Spec Maker](../../skills/implementation-spec-maker/SKILL.md)
|
|
5
|
-
for the supplied issue authority. Use the declared
|
|
6
|
-
[confidence rubric](../../docs/agents/confidence-rubric.md) and
|
|
7
|
-
[contract ledger](../../docs/agents/contract-test-ledger.md) only when
|
|
8
|
-
applicable.
|
|
9
|
-
The Runner owns artifact review, revision state, and approval, so do not invoke
|
|
10
|
-
the skill's review/save workflow or create reviewer state. Write the complete
|
|
11
|
-
revision only to the Runner-provided spec artifact location and return
|
|
12
|
-
`schemas/spec-author-v1.json`.
|
|
@@ -1,12 +0,0 @@
|
|
|
1
|
-
# Spec Review Operation
|
|
2
|
-
|
|
3
|
-
You are already the independent reviewer selected and persisted by the Runner.
|
|
4
|
-
Follow the packaged
|
|
5
|
-
[Implementation Spec Review](../../skills/implementation-spec-review/SKILL.md)
|
|
6
|
-
inline and use its owner-local review loop only for semantics matching the
|
|
7
|
-
supplied mode and immutable state. Apply the declared
|
|
8
|
-
[confidence rubric](../../docs/agents/confidence-rubric.md),
|
|
9
|
-
[contract ledger](../../docs/agents/contract-test-ledger.md), and
|
|
10
|
-
[review protocol](../../docs/agents/review-protocol.md). Do not launch another
|
|
11
|
-
reviewer, edit the spec, change review state, or mutate external state. Return
|
|
12
|
-
only `schemas/spec-review-v1.json`.
|