codex-orchestrator 2.0.1 → 2.0.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +22 -0
- package/README.md +12 -9
- package/dist/src/index.d.ts +10 -0
- package/dist/src/index.d.ts.map +1 -1
- package/dist/src/index.js +5 -0
- package/dist/src/index.js.map +1 -1
- package/dist/src/v2/acceptance-proof.d.ts +3 -0
- package/dist/src/v2/acceptance-proof.d.ts.map +1 -1
- package/dist/src/v2/acceptance-proof.js +2 -8
- package/dist/src/v2/acceptance-proof.js.map +1 -1
- package/dist/src/v2/adapters/gh-issue-adapter.d.ts +5 -3
- package/dist/src/v2/adapters/gh-issue-adapter.d.ts.map +1 -1
- package/dist/src/v2/adapters/gh-issue-adapter.js +63 -7
- package/dist/src/v2/adapters/gh-issue-adapter.js.map +1 -1
- package/dist/src/v2/adapters/issues.d.ts +16 -2
- package/dist/src/v2/adapters/issues.d.ts.map +1 -1
- package/dist/src/v2/adapters/issues.js +15 -5
- package/dist/src/v2/adapters/issues.js.map +1 -1
- package/dist/src/v2/adapters/mission-coordinator-lock.d.ts +1 -0
- package/dist/src/v2/adapters/mission-coordinator-lock.d.ts.map +1 -1
- package/dist/src/v2/adapters/mission-coordinator-lock.js +5 -1
- package/dist/src/v2/adapters/mission-coordinator-lock.js.map +1 -1
- package/dist/src/v2/candidate-cli.d.ts +4 -0
- package/dist/src/v2/candidate-cli.d.ts.map +1 -1
- package/dist/src/v2/candidate-cli.js +26 -11
- package/dist/src/v2/candidate-cli.js.map +1 -1
- package/dist/src/v2/cli-contract.d.ts +1 -1
- package/dist/src/v2/cli-contract.d.ts.map +1 -1
- package/dist/src/v2/cli-contract.js +10 -0
- package/dist/src/v2/cli-contract.js.map +1 -1
- package/dist/src/v2/code-review-report.d.ts +66 -0
- package/dist/src/v2/code-review-report.d.ts.map +1 -0
- package/dist/src/v2/code-review-report.js +259 -0
- package/dist/src/v2/code-review-report.js.map +1 -0
- package/dist/src/v2/codex-process.d.ts +8 -1
- package/dist/src/v2/codex-process.d.ts.map +1 -1
- package/dist/src/v2/codex-process.js +11 -0
- package/dist/src/v2/codex-process.js.map +1 -1
- package/dist/src/v2/config.d.ts +2 -1
- package/dist/src/v2/config.d.ts.map +1 -1
- package/dist/src/v2/config.js +8 -3
- package/dist/src/v2/config.js.map +1 -1
- package/dist/src/v2/contained-report-operation.d.ts +100 -0
- package/dist/src/v2/contained-report-operation.d.ts.map +1 -0
- package/dist/src/v2/contained-report-operation.js +200 -0
- package/dist/src/v2/contained-report-operation.js.map +1 -0
- package/dist/src/v2/containment.d.ts +6 -0
- package/dist/src/v2/containment.d.ts.map +1 -1
- package/dist/src/v2/containment.js +40 -1
- package/dist/src/v2/containment.js.map +1 -1
- package/dist/src/v2/direct-delivery.d.ts +101 -0
- package/dist/src/v2/direct-delivery.d.ts.map +1 -0
- package/dist/src/v2/direct-delivery.js +547 -0
- package/dist/src/v2/direct-delivery.js.map +1 -0
- package/dist/src/v2/immutable-workflow-publisher.d.ts +40 -0
- package/dist/src/v2/immutable-workflow-publisher.d.ts.map +1 -0
- package/dist/src/v2/immutable-workflow-publisher.js +218 -0
- package/dist/src/v2/immutable-workflow-publisher.js.map +1 -0
- package/dist/src/v2/implementation-reviewer.d.ts +81 -0
- package/dist/src/v2/implementation-reviewer.d.ts.map +1 -0
- package/dist/src/v2/implementation-reviewer.js +157 -0
- package/dist/src/v2/implementation-reviewer.js.map +1 -0
- package/dist/src/v2/owner-control-lock.d.ts +41 -0
- package/dist/src/v2/owner-control-lock.d.ts.map +1 -0
- package/dist/src/v2/owner-control-lock.js +174 -0
- package/dist/src/v2/owner-control-lock.js.map +1 -0
- package/dist/src/v2/route-continuations.d.ts +32 -0
- package/dist/src/v2/route-continuations.d.ts.map +1 -0
- package/dist/src/v2/route-continuations.js +2 -0
- package/dist/src/v2/route-continuations.js.map +1 -0
- package/dist/src/v2/route-coordinator.d.ts +77 -0
- package/dist/src/v2/route-coordinator.d.ts.map +1 -0
- package/dist/src/v2/route-coordinator.js +370 -0
- package/dist/src/v2/route-coordinator.js.map +1 -0
- package/dist/src/v2/route-decision.d.ts +129 -0
- package/dist/src/v2/route-decision.d.ts.map +1 -0
- package/dist/src/v2/route-decision.js +400 -0
- package/dist/src/v2/route-decision.js.map +1 -0
- package/dist/src/v2/run-issue.d.ts +63 -2
- package/dist/src/v2/run-issue.d.ts.map +1 -1
- package/dist/src/v2/run-issue.js +906 -91
- package/dist/src/v2/run-issue.js.map +1 -1
- package/dist/src/v2/run-store.d.ts +25 -1
- package/dist/src/v2/run-store.d.ts.map +1 -1
- package/dist/src/v2/run-store.js +143 -3
- package/dist/src/v2/run-store.js.map +1 -1
- package/dist/src/v2/runtime-assets.d.ts +15 -13
- package/dist/src/v2/runtime-assets.d.ts.map +1 -1
- package/dist/src/v2/runtime-assets.js +263 -416
- package/dist/src/v2/runtime-assets.js.map +1 -1
- package/dist/src/v2/runtime.d.ts +14 -6
- package/dist/src/v2/runtime.d.ts.map +1 -1
- package/dist/src/v2/runtime.js +478 -56
- package/dist/src/v2/runtime.js.map +1 -1
- package/dist/src/v2/setup-cli.d.ts.map +1 -1
- package/dist/src/v2/setup-cli.js +1 -0
- package/dist/src/v2/setup-cli.js.map +1 -1
- package/dist/src/v2/setup-runtime.d.ts.map +1 -1
- package/dist/src/v2/setup-runtime.js +20 -72
- package/dist/src/v2/setup-runtime.js.map +1 -1
- package/dist/src/v2/setup.d.ts +4 -1
- package/dist/src/v2/setup.d.ts.map +1 -1
- package/dist/src/v2/setup.js +104 -1
- package/dist/src/v2/setup.js.map +1 -1
- package/dist/src/v2/spec-coordinator.d.ts +85 -0
- package/dist/src/v2/spec-coordinator.d.ts.map +1 -0
- package/dist/src/v2/spec-coordinator.js +88 -0
- package/dist/src/v2/spec-coordinator.js.map +1 -0
- package/dist/src/v2/spec-delivery.d.ts +143 -0
- package/dist/src/v2/spec-delivery.d.ts.map +1 -0
- package/dist/src/v2/spec-delivery.js +401 -0
- package/dist/src/v2/spec-delivery.js.map +1 -0
- package/dist/src/v2/triage-route.d.ts +68 -0
- package/dist/src/v2/triage-route.d.ts.map +1 -0
- package/dist/src/v2/triage-route.js +223 -0
- package/dist/src/v2/triage-route.js.map +1 -0
- package/dist/src/v2/waiting-human-coordinator.d.ts +49 -0
- package/dist/src/v2/waiting-human-coordinator.d.ts.map +1 -0
- package/dist/src/v2/waiting-human-coordinator.js +509 -0
- package/dist/src/v2/waiting-human-coordinator.js.map +1 -0
- package/dist/src/v2/waiting-human.d.ts +143 -0
- package/dist/src/v2/waiting-human.d.ts.map +1 -0
- package/dist/src/v2/waiting-human.js +408 -0
- package/dist/src/v2/waiting-human.js.map +1 -0
- package/dist/src/v2/workflow-assets.d.ts +90 -0
- package/dist/src/v2/workflow-assets.d.ts.map +1 -0
- package/dist/src/v2/workflow-assets.js +554 -0
- package/dist/src/v2/workflow-assets.js.map +1 -0
- package/docs/deep-dive.md +15 -8
- package/internal-workflow/docs/agents/artifact-review-loop.md +267 -0
- package/internal-workflow/docs/agents/bug-workflow-routing.md +24 -0
- package/internal-workflow/docs/agents/coding-skill-routing.md +203 -0
- package/internal-workflow/docs/agents/confidence-rubric.md +65 -0
- package/internal-workflow/docs/agents/contract-test-ledger.md +60 -0
- package/internal-workflow/docs/agents/implementation-review-loop.md +302 -0
- package/internal-workflow/docs/agents/review-gates.md +49 -0
- package/internal-workflow/docs/agents/review-protocol.md +170 -0
- package/internal-workflow/docs/agents/tool-usage.md +88 -0
- package/internal-workflow/manifest.json +1 -0
- package/internal-workflow/operations/acceptance-proof/SKILL.md +3 -0
- package/internal-workflow/operations/ambiguity-review/SKILL.md +3 -0
- package/internal-workflow/operations/cleanup-review/SKILL.md +3 -0
- package/internal-workflow/operations/code-review/SKILL.md +3 -0
- package/internal-workflow/operations/implementation/SKILL.md +3 -0
- package/internal-workflow/operations/spec-author/SKILL.md +3 -0
- package/internal-workflow/operations/spec-implementation/SKILL.md +3 -0
- package/internal-workflow/operations/spec-review/SKILL.md +3 -0
- package/internal-workflow/operations/triage/SKILL.md +3 -0
- package/internal-workflow/profiles/analyst_deep.toml +9 -0
- package/internal-workflow/profiles/implementer_deep.toml +9 -0
- package/internal-workflow/profiles/implementer_standard.toml +9 -0
- package/internal-workflow/profiles/proof_agent.toml +8 -0
- package/internal-workflow/profiles/researcher_standard.toml +9 -0
- package/internal-workflow/profiles/reviewer_deep.toml +9 -0
- package/internal-workflow/profiles/reviewer_fast.toml +9 -0
- package/internal-workflow/profiles/reviewer_standard.toml +9 -0
- package/internal-workflow/schemas/ambiguity-review-v1.json +1 -0
- package/internal-workflow/schemas/code-review-v1.json +1 -0
- package/internal-workflow/schemas/implementation-report-v1.json +1 -0
- package/internal-workflow/schemas/proof-report-v1.json +1 -0
- package/internal-workflow/schemas/spec-author-v1.json +1 -0
- package/internal-workflow/schemas/spec-review-v1.json +30 -0
- package/internal-workflow/schemas/triage-route-v1.json +1 -0
- package/internal-workflow/skills/acceptance-proof/agents/openai.yaml +6 -0
- package/internal-workflow/skills/agent-auto/agents/openai.yaml +6 -0
- package/internal-workflow/skills/cleanup-review/SKILL.md +84 -0
- package/internal-workflow/skills/cleanup-review/agents/openai.yaml +6 -0
- package/internal-workflow/skills/code-review/SKILL.md +257 -0
- package/internal-workflow/skills/code-review/agents/openai.yaml +4 -0
- package/internal-workflow/skills/code-review/references/bug-classes.md +56 -0
- package/internal-workflow/skills/code-review/references/framework-lenses.md +34 -0
- package/internal-workflow/skills/code-review/references/targeted-recipes.md +49 -0
- package/internal-workflow/skills/codebase-design/DEEPENING.md +35 -0
- package/internal-workflow/skills/codebase-design/DESIGN-IT-TWICE.md +50 -0
- package/internal-workflow/skills/codebase-design/SKILL.md +82 -0
- package/internal-workflow/skills/codebase-design/agents/openai.yaml +6 -0
- package/internal-workflow/skills/diagnosing-bugs/SKILL.md +138 -0
- package/internal-workflow/skills/diagnosing-bugs/agents/openai.yaml +6 -0
- package/internal-workflow/skills/diagnosing-bugs/scripts/hitl-loop.template.sh +41 -0
- package/internal-workflow/skills/implementation-spec-maker/SKILL.md +93 -0
- package/internal-workflow/skills/implementation-spec-maker/agents/openai.yaml +6 -0
- package/internal-workflow/skills/implementation-spec-maker/references/source-modes.md +31 -0
- package/internal-workflow/skills/implementation-spec-maker/references/spec-template.md +146 -0
- package/internal-workflow/skills/implementation-spec-review/SKILL.md +211 -0
- package/internal-workflow/skills/implementation-spec-review/agents/openai.yaml +6 -0
- package/internal-workflow/skills/research/SKILL.md +107 -0
- package/internal-workflow/skills/research/agents/openai.yaml +6 -0
- package/internal-workflow/skills/small-task-implementer/SKILL.md +97 -0
- package/internal-workflow/skills/small-task-implementer/agents/openai.yaml +6 -0
- package/internal-workflow/skills/spec-implementer/SKILL.md +197 -0
- package/internal-workflow/skills/spec-implementer/agents/openai.yaml +6 -0
- package/internal-workflow/skills/tdd/SKILL.md +59 -0
- package/internal-workflow/skills/tdd/agents/openai.yaml +6 -0
- package/internal-workflow/skills/tdd/interface-design.md +31 -0
- package/internal-workflow/skills/tdd/mocking.md +59 -0
- package/internal-workflow/skills/tdd/refactoring.md +10 -0
- package/internal-workflow/skills/tdd/tests.md +77 -0
- package/internal-workflow/skills/triage/AGENT-BRIEF.md +192 -0
- package/internal-workflow/skills/triage/OUT-OF-SCOPE.md +101 -0
- package/internal-workflow/skills/triage/SKILL.md +134 -0
- package/internal-workflow/skills/triage/agents/openai.yaml +6 -0
- package/internal-workflow/skills/ui-evidence-proof/SKILL.md +123 -0
- package/internal-workflow/skills/ui-evidence-proof/agents/openai.yaml +6 -0
- package/package.json +6 -3
- /package/{internal-skills → internal-workflow/skills}/acceptance-proof/SKILL.md +0 -0
- /package/{internal-skills → internal-workflow/skills}/acceptance-proof/references/android.md +0 -0
- /package/{internal-skills → internal-workflow/skills}/acceptance-proof/references/browser.md +0 -0
- /package/{internal-skills → internal-workflow/skills}/acceptance-proof/references/ios.md +0 -0
- /package/{internal-skills → internal-workflow/skills}/acceptance-proof/tools/android-lease.mjs +0 -0
- /package/{internal-skills → internal-workflow/skills}/acceptance-proof/tools/ios-lease.mjs +0 -0
- /package/{internal-skills → internal-workflow/skills}/agent-auto/SKILL.md +0 -0
|
@@ -0,0 +1,65 @@
|
|
|
1
|
+
# Confidence Rubric
|
|
2
|
+
|
|
3
|
+
Use this rubric for coding skills that report findings, diagnose root causes, review specs, or decide whether an auto-fix is safe.
|
|
4
|
+
|
|
5
|
+
Do not use fake numeric precision unless a skill has a concrete scoring reason. Prefer `high`, `medium`, and `low` with evidence.
|
|
6
|
+
|
|
7
|
+
## High Confidence
|
|
8
|
+
|
|
9
|
+
High confidence means the finding or diagnosis has direct evidence.
|
|
10
|
+
|
|
11
|
+
Requires:
|
|
12
|
+
|
|
13
|
+
- a concrete trigger path, code path, failing signal, or tool output
|
|
14
|
+
- a clear explanation of why existing guards do not prevent the issue
|
|
15
|
+
- no unresolved assumption that changes the conclusion
|
|
16
|
+
|
|
17
|
+
Allowed actions:
|
|
18
|
+
|
|
19
|
+
- report as a finding
|
|
20
|
+
- block execution when the issue is safety-critical or spec-critical
|
|
21
|
+
- auto-fix only when the fix is narrow, low-risk, local-patterned, and verifiable
|
|
22
|
+
|
|
23
|
+
## Medium Confidence
|
|
24
|
+
|
|
25
|
+
Medium confidence means the issue is likely but one explicit assumption remains.
|
|
26
|
+
|
|
27
|
+
Requires:
|
|
28
|
+
|
|
29
|
+
- strong local evidence
|
|
30
|
+
- exactly what assumption remains
|
|
31
|
+
- what evidence would promote or demote the finding
|
|
32
|
+
|
|
33
|
+
Allowed actions:
|
|
34
|
+
|
|
35
|
+
- report as a likely issue, risk, or execution concern
|
|
36
|
+
- ask a targeted question when the unresolved assumption changes the fix
|
|
37
|
+
- do not auto-fix unless new evidence raises confidence to high
|
|
38
|
+
|
|
39
|
+
## Low Confidence
|
|
40
|
+
|
|
41
|
+
Low confidence means the concern is plausible but not proven.
|
|
42
|
+
|
|
43
|
+
Requires:
|
|
44
|
+
|
|
45
|
+
- a clear label as uncertainty
|
|
46
|
+
- the missing evidence or verification gap
|
|
47
|
+
|
|
48
|
+
Allowed actions:
|
|
49
|
+
|
|
50
|
+
- present as a question, risk, or verification gap
|
|
51
|
+
- do not report as a proven bug
|
|
52
|
+
- do not auto-fix
|
|
53
|
+
|
|
54
|
+
## Auto-Fix Gate
|
|
55
|
+
|
|
56
|
+
Auto-fix is allowed only when all are true:
|
|
57
|
+
|
|
58
|
+
- confidence is high
|
|
59
|
+
- root cause is clear
|
|
60
|
+
- fix is narrow and low-risk
|
|
61
|
+
- fix matches local project patterns
|
|
62
|
+
- verification is available, or the edit is syntax-checkable and obviously safe
|
|
63
|
+
- the change does not require a product decision
|
|
64
|
+
|
|
65
|
+
If any condition is missing, report the issue with evidence and stop before editing.
|
|
@@ -0,0 +1,60 @@
|
|
|
1
|
+
# Contract Test Ledger
|
|
2
|
+
|
|
3
|
+
Use this shared ledger for behavior-changing work where a passing happy-path test could still miss a contract defect. The ledger turns review-class risks into testable obligations before implementation.
|
|
4
|
+
|
|
5
|
+
## When Required
|
|
6
|
+
|
|
7
|
+
Create or update a contract test ledger when the task changes any of these:
|
|
8
|
+
|
|
9
|
+
- API, DTO, schema, serialization, persistence, or externally visible response shape
|
|
10
|
+
- ordering, lifecycle events, state transitions, retries, idempotency, timeout, cancellation, or background jobs
|
|
11
|
+
- cache keys, invalidation, state merge precedence, fallback behavior, profile/global/mobile overrides, defaults, or feature flags
|
|
12
|
+
- evidence, trace, snapshot, audit, summary, aggregation, score, winner, or generated artifacts
|
|
13
|
+
- shared behavior read by multiple callers, tenants, users, groups, children, or projections
|
|
14
|
+
|
|
15
|
+
For narrow UI copy, docs-only, formatting, tests-only, or isolated styling changes, the ledger is not required.
|
|
16
|
+
|
|
17
|
+
## Required Shape
|
|
18
|
+
|
|
19
|
+
Keep the ledger compact. Use one row per invariant:
|
|
20
|
+
|
|
21
|
+
```markdown
|
|
22
|
+
## Contract Test Ledger
|
|
23
|
+
|
|
24
|
+
| Invariant | Risk It Prevents | First Test / Proof | Status |
|
|
25
|
+
| --- | --- | --- | --- |
|
|
26
|
+
| <observable rule> | <real failure mode> | <exact test name/command or manual proof> | planned / red / green / blocked |
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
Rules:
|
|
30
|
+
|
|
31
|
+
- The invariant must be observable through the public interface or the same seam real callers use.
|
|
32
|
+
- The risk must name the concrete bug class, not a vague "edge case".
|
|
33
|
+
- The first test/proof must fail before the fix unless the ledger records why a RED signal is impossible.
|
|
34
|
+
- `blocked` requires the missing seam, fixture, service, or decision that prevents proof.
|
|
35
|
+
- Keep the ledger current as implementation proceeds; do not backfill it only at the end.
|
|
36
|
+
|
|
37
|
+
## Invariant Prompts
|
|
38
|
+
|
|
39
|
+
Ask the relevant subset before the first RED test:
|
|
40
|
+
|
|
41
|
+
- **Ordering:** What must happen before/after terminal events, snapshots, persistence writes, notifications, or cleanup?
|
|
42
|
+
- **Precedence:** Which source wins among user input, profile, mobile, global, server, cache, AI, fallback, default, `null`, `false`, `0`, and empty objects?
|
|
43
|
+
- **Threading:** Does each new field survive construction, normalization, cloning, retry, persistence reload, serialization, and every visible consumer?
|
|
44
|
+
- **Runtime contract:** Do validation, internal types, persistence schema, API response, and consumers agree on names, units, nullability, enum values, and date/object formats?
|
|
45
|
+
- **Retry/idempotency:** What changes on retry, and which snapshots, counters, streams, timestamps, writes, or side effects must be rebuilt instead of reused?
|
|
46
|
+
- **Determinism:** When sort keys, timestamps, scores, priorities, or winners tie, what stable tie-breaker makes output repeatable?
|
|
47
|
+
- **Evidence:** Which trace, audit, snapshot, Fresh-Context, summary, or generated artifact proves the behavior actually happened?
|
|
48
|
+
- **Partial failure:** If a dependency times out, throws, returns stale data, or fails after a side effect, what durable state remains and who repairs it?
|
|
49
|
+
- **Scope/cardinality:** Is data global, per-tenant, per-group, per-child, per-step, or per-item, and can one top-level field collapse multiple meaningful results?
|
|
50
|
+
|
|
51
|
+
## Review Feedback Loop
|
|
52
|
+
|
|
53
|
+
When code review finds a real contract defect, add or update one ledger row before fixing it:
|
|
54
|
+
|
|
55
|
+
- `Invariant`: the rule the implementation violated
|
|
56
|
+
- `Risk It Prevents`: the observed review finding
|
|
57
|
+
- `First Test / Proof`: the regression test or proof that would have caught it
|
|
58
|
+
- `Status`: `red` before the fix, then `green` after verification
|
|
59
|
+
|
|
60
|
+
If no correct public seam exists for the regression test, record that as `blocked` and name the architecture/testability gap. Do not replace a missing seam with an implementation-detail test unless the task explicitly approves that tradeoff.
|
|
@@ -0,0 +1,302 @@
|
|
|
1
|
+
# Implementation Review Loop
|
|
2
|
+
|
|
3
|
+
Use this policy for every approved implementation-spec execution. This
|
|
4
|
+
target-specific Module owns implementation authority, durable Review State,
|
|
5
|
+
checkpoint/final topology, validation reuse, gate ordering, audit epochs, and
|
|
6
|
+
implementation outcome mapping. It applies
|
|
7
|
+
[`review-protocol.md`](review-protocol.md) for context transfer, Full/Closure
|
|
8
|
+
mechanics, defect lifecycle, no-progress, and the common result envelope.
|
|
9
|
+
`spec-implementer`, `cleanup-review`, and `code-review` are callers or Adapters;
|
|
10
|
+
they must not reproduce either Module. Deterministic tickets-orchestrator work
|
|
11
|
+
using its issue as authority remains outside this Module and uses direct TDD
|
|
12
|
+
plus repo review gates.
|
|
13
|
+
|
|
14
|
+
This policy does not replace tests, architecture checks, smoke tests, or Git
|
|
15
|
+
checkpoints. Those proofs remain independent evidence.
|
|
16
|
+
|
|
17
|
+
## Interface
|
|
18
|
+
|
|
19
|
+
Conceptually, the executor uses:
|
|
20
|
+
|
|
21
|
+
```text
|
|
22
|
+
review_implementation(
|
|
23
|
+
authority_artifact_kind,
|
|
24
|
+
authority_artifact_path,
|
|
25
|
+
review_profile,
|
|
26
|
+
current_revision,
|
|
27
|
+
checkpoint,
|
|
28
|
+
review_focus,
|
|
29
|
+
defect_ledger
|
|
30
|
+
) -> review_outcome
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
The outcome records:
|
|
34
|
+
|
|
35
|
+
- `outcome`: `Approved | Blocked | Waived`
|
|
36
|
+
- `authority_artifact_kind`: `approved-spec`
|
|
37
|
+
- `authority_artifact_path`: the sole artifact that stores review state
|
|
38
|
+
- `review_profile`: `simple | medium | high`
|
|
39
|
+
- completed review passes and pending launches
|
|
40
|
+
- review mode, reviewer/session identity, target revision, and assigned lenses
|
|
41
|
+
- stable defect IDs and their current status
|
|
42
|
+
- logical skill activations and their open/closed state
|
|
43
|
+
- mandatory final coverage still required
|
|
44
|
+
|
|
45
|
+
## Authority
|
|
46
|
+
|
|
47
|
+
An approved implementation spec stores the sole `## Implementation Review
|
|
48
|
+
State`. Architecture RFCs, product PRDs, tickets, `ready-for-approval`
|
|
49
|
+
artifacts, inferred approval, and ordinary direct-ticket waves are ineligible.
|
|
50
|
+
Persist `authority_artifact_kind` and `authority_artifact_path` and never create
|
|
51
|
+
a second ledger in an upstream artifact, caller, ticket, or Adapter.
|
|
52
|
+
|
|
53
|
+
Lifecycle and proof updates to the selected artifact do not change its approved
|
|
54
|
+
status or substantive design. A substantive authority change requires the
|
|
55
|
+
normal artifact revision/review path before implementation continues.
|
|
56
|
+
|
|
57
|
+
## Profile And Review Shape
|
|
58
|
+
|
|
59
|
+
Prefer the selected authority artifact's `review_profile`. If it is absent, use the same
|
|
60
|
+
evidence-based classification and hard escalators as
|
|
61
|
+
[`artifact-review-loop.md`](artifact-review-loop.md). Actual implementation
|
|
62
|
+
evidence may raise the profile but must not lower it.
|
|
63
|
+
|
|
64
|
+
Review profile selects mandatory lenses and independence. Protocol pass-count
|
|
65
|
+
semantics apply; parallel reviewer results remain separate passes even when they
|
|
66
|
+
reduce wall-clock time.
|
|
67
|
+
|
|
68
|
+
Select the reviewer role from the profile: `simple` uses `reviewer_fast`,
|
|
69
|
+
`medium` uses `reviewer_standard`, and `high` uses `reviewer_deep`. Root always
|
|
70
|
+
launches reviewer children; it never performs an implementation review inline.
|
|
71
|
+
An Adapter runs inline only inside its already assigned reviewer child.
|
|
72
|
+
|
|
73
|
+
A reviewer that fails before returning a usable result is recorded as failed
|
|
74
|
+
and closed, not as completed coverage. Escalation preserves completed coverage
|
|
75
|
+
and the stable Defect Ledger. A user pause, context compaction, slice commit,
|
|
76
|
+
worker replacement, or new turn also preserves them; none restarts the review
|
|
77
|
+
topology automatically.
|
|
78
|
+
|
|
79
|
+
Stop early when mandatory coverage and protocol clear-state requirements hold.
|
|
80
|
+
|
|
81
|
+
## Review Planning
|
|
82
|
+
|
|
83
|
+
Before the first implementation reviewer, root creates a short Review Plan:
|
|
84
|
+
|
|
85
|
+
- current profile and required independent lenses
|
|
86
|
+
- only stable intermediate checkpoints and their required lenses; move an unstable checkpoint to final coverage when later slices touch the same files, owners, or contracts
|
|
87
|
+
- any separate cleanup requirement, which must name a concrete evidenced reason that cannot fit the final spec/standards lens
|
|
88
|
+
- final code-review lenses and minimum independent coverage
|
|
89
|
+
- reviewer lineages that own affected-lens Closure
|
|
90
|
+
|
|
91
|
+
Create one durable activation record for each logical skill invocation. Record
|
|
92
|
+
`activation_id`, skill, owner, opened/closed state, and resume rule. Review Full
|
|
93
|
+
and lineage-preserving Closure passes stay inside that review skill's activation;
|
|
94
|
+
TDD repair cycles stay inside the active TDD activation. Cleanup, code review,
|
|
95
|
+
TDD, and debugger activations never share an ID, and a continuation resumes an
|
|
96
|
+
ID only for the same skill and authorized flow.
|
|
97
|
+
|
|
98
|
+
## Durable Review State
|
|
99
|
+
|
|
100
|
+
Do not create durable review state during implementation preflight. Immediately
|
|
101
|
+
before the first actual reviewer launch, persist the short Review Plan and
|
|
102
|
+
pending launch in the selected authority artifact under
|
|
103
|
+
`## Implementation Review State`. From that point onward this is the execution
|
|
104
|
+
ledger for review state; do not keep the authoritative history only in chat
|
|
105
|
+
context or a subagent summary.
|
|
106
|
+
|
|
107
|
+
Record at least:
|
|
108
|
+
|
|
109
|
+
- profile, completed pass count, and required coverage still outstanding
|
|
110
|
+
- current checkpoint/gate, review timing baseline, gate-local consecutive
|
|
111
|
+
Closure-wave count, and latest Closure wave ID
|
|
112
|
+
- planned mandatory final reviews and their lenses
|
|
113
|
+
- authority artifact kind/path and logical skill activation records
|
|
114
|
+
- each lineage ID, origin Full session, active session generation, Closure count,
|
|
115
|
+
rotation reason, live/timeout state, and `conclude_requested_at`
|
|
116
|
+
- any convergence audit epoch: trigger, triggering pass/wave/revision,
|
|
117
|
+
completion, dispositions, selected sessions, resume reason, and pass/wave/time
|
|
118
|
+
baselines used for its next rearm
|
|
119
|
+
- pending reviewer launches with launch ID, mode, lineage/session identity,
|
|
120
|
+
activation ID, target revision, checkpoint/gate, Closure wave ID when
|
|
121
|
+
applicable, assigned lenses, and start timestamp
|
|
122
|
+
- every completed review's mode, lineage/session identity, target revision,
|
|
123
|
+
checkpoint/gate, Closure wave ID, assigned lenses, start/end timestamps, and
|
|
124
|
+
outcome
|
|
125
|
+
- the stable Defect Ledger with transition history, reopen count, fixed revision,
|
|
126
|
+
verifying review, and any explicit risk acceptance
|
|
127
|
+
|
|
128
|
+
Write a pending launch before starting its reviewer. After the launch returns,
|
|
129
|
+
replace the pending record with either its usable completed result or a failed,
|
|
130
|
+
closed session record; only a usable result increments `review_passes`. A context
|
|
131
|
+
compaction, new turn, resumed task, or different root agent must reconcile every
|
|
132
|
+
pending launch with its recorded session before starting a replacement.
|
|
133
|
+
|
|
134
|
+
Update this section after each launch, usable reviewer result, repair batch,
|
|
135
|
+
closure, waiver, acceptance, reopen, or terminal outcome. A resumed executor
|
|
136
|
+
reconstructs accounting and lifecycle history from this persisted state. If the
|
|
137
|
+
state is missing or internally inconsistent after reviews began, return
|
|
138
|
+
`Blocked` until it is reconciled from available thread/session evidence; never
|
|
139
|
+
assume zero completed passes or silently replace an in-flight reviewer.
|
|
140
|
+
|
|
141
|
+
Plan mandatory final coverage before launching an intermediate review. Launch a
|
|
142
|
+
checkpoint only when its target is settled and later slices will not invalidate
|
|
143
|
+
the reviewed files, owners, or contracts. Otherwise move its lenses to final
|
|
144
|
+
coverage. Do not replace a required final lens with another fresh checkpoint
|
|
145
|
+
reviewer or a repeat broad audit.
|
|
146
|
+
|
|
147
|
+
The default shapes are:
|
|
148
|
+
|
|
149
|
+
- `simple`: validation only when policy does not require review; otherwise one
|
|
150
|
+
final Full review and affected-lens Closure only after repairs.
|
|
151
|
+
- `medium`: an explicit intermediate checkpoint may provide one required lens;
|
|
152
|
+
the final integrator covers every remaining lens, includes bounded cleanup in
|
|
153
|
+
spec/standards, and verifies its defects without a separate cleanup pass.
|
|
154
|
+
- `high`: use parallel independent tracks only for disjoint mandatory lenses;
|
|
155
|
+
the spec/standards track includes bounded cleanup, and affected-lens Closure
|
|
156
|
+
follows only after consolidated repairs. There is no separate cleanup pass by
|
|
157
|
+
default.
|
|
158
|
+
|
|
159
|
+
`code-review` uses one final reviewer covering both correctness and
|
|
160
|
+
spec/standards for `simple` and `medium`. For `high`, it uses two disjoint final
|
|
161
|
+
tracks unless earlier independent coverage already covered both axes and one
|
|
162
|
+
fresh final integrator receives their compact handoffs.
|
|
163
|
+
|
|
164
|
+
Before launching a fresh final reviewer, reconcile coverage on the settled
|
|
165
|
+
revision. If the latest usable Full or Closure covered every mandatory final
|
|
166
|
+
lens and left no open defect, mark final review complete and stop. Count cleanup
|
|
167
|
+
Closure only for the lenses explicitly assigned in the Review Plan; a
|
|
168
|
+
`cleanup-only` pass does not satisfy correctness or spec/standards coverage.
|
|
169
|
+
|
|
170
|
+
## Review Capsule
|
|
171
|
+
|
|
172
|
+
Use the protocol capsule with these implementation fields:
|
|
173
|
+
|
|
174
|
+
- the unanswered implementation question and any prior coverage it invalidates
|
|
175
|
+
- authority artifact kind/path, profile, current revision, checkpoint, and exact diff command
|
|
176
|
+
- changed paths and assigned `Review Focus` lenses
|
|
177
|
+
- source-of-truth docs and relevant Contract Test Ledger rows
|
|
178
|
+
- compact validation results and known verification gaps
|
|
179
|
+
|
|
180
|
+
For Closure, map changed paths and tests into the protocol Revision Map.
|
|
181
|
+
|
|
182
|
+
## Review Modes
|
|
183
|
+
|
|
184
|
+
Use protocol Full and Closure without redefining them. Implementation Closure
|
|
185
|
+
maps `affected_targets` to paths, tests, runtime contracts, and Review Focus
|
|
186
|
+
lenses. An already planned Full reviewer may verify a repair when its assigned
|
|
187
|
+
lenses cover it.
|
|
188
|
+
|
|
189
|
+
## Defect Lifecycle
|
|
190
|
+
|
|
191
|
+
Use the canonical protocol ledger and lifecycle without local aliases.
|
|
192
|
+
|
|
193
|
+
Implementation proof-only gaps may use `planned-final-verification` only when an
|
|
194
|
+
already scheduled code-review lens owns the proof; they remain open until
|
|
195
|
+
independently verified and never re-enter cleanup. Artifact proof-contract gaps
|
|
196
|
+
reopen artifact review. A pre-existing adjacent issue is non-blocking only as an
|
|
197
|
+
`improvement` with `follow-up-improvement`.
|
|
198
|
+
|
|
199
|
+
## Validation Evidence Reuse
|
|
200
|
+
|
|
201
|
+
Persist command/config identity, failure signature, target revision and changed
|
|
202
|
+
path/contract impact basis, secret-safe environment fingerprint, transitive
|
|
203
|
+
ownership/contract impact, and result. Reuse a known unrelated suite failure
|
|
204
|
+
only when every field matches; unknown environment or transitive impact fails
|
|
205
|
+
closed. Focused tests and every required check for a repair always rerun.
|
|
206
|
+
|
|
207
|
+
## Gate Ordering
|
|
208
|
+
|
|
209
|
+
1. Implement the slice and pass its tests/exit gate.
|
|
210
|
+
2. At an explicit intermediate checkpoint, run the required targeted
|
|
211
|
+
`code-review` directly under the Review Plan.
|
|
212
|
+
3. Repair one consolidated finding batch and use protocol Closure for the
|
|
213
|
+
affected lineages. An already planned Full reviewer may verify the repair when
|
|
214
|
+
its assigned lenses cover it.
|
|
215
|
+
At one gate, collect the usable results from all already-launched reviewers
|
|
216
|
+
before repairing, unless an immediate blocker invalidates the remaining
|
|
217
|
+
work. Do not turn individual findings into serial repair, validation, and
|
|
218
|
+
Closure micro-cycles. Repair compatible findings once, rerun each affected
|
|
219
|
+
validation once on the resulting revision, then launch one affected-lens
|
|
220
|
+
Closure wave.
|
|
221
|
+
4. Continue implementation only when checkpoint blockers are verified or the
|
|
222
|
+
Review Plan explicitly assigns their verification to an already planned reviewer
|
|
223
|
+
without violating the checkpoint's safety purpose.
|
|
224
|
+
5. After all implementation slices and validations settle, run one final code
|
|
225
|
+
review wave. `simple` and `medium` use one reviewer; `high` launches two
|
|
226
|
+
disjoint reviewer tracks in parallel. The spec/standards lens owns bounded
|
|
227
|
+
cleanup.
|
|
228
|
+
|
|
229
|
+
Intermediate code-review checkpoints do not run cleanup-review. A separate
|
|
230
|
+
cleanup pass is exceptional: run it only when the user, approved source, or repo
|
|
231
|
+
policy names a concrete evidenced simplification risk that cannot fit the final
|
|
232
|
+
spec/standards lens. `large` or `high` alone is not a reason. If an approved spec
|
|
233
|
+
names an intermediate cleanup checkpoint, return `Blocked` for spec revision.
|
|
234
|
+
|
|
235
|
+
Cleanup review runs at most once as a Full review for the whole spec. After its
|
|
236
|
+
findings are repaired, either use protocol Closure or give the final code
|
|
237
|
+
reviewer those stable defect IDs for verification. Never launch another Full
|
|
238
|
+
cleanup review over the repaired whole diff.
|
|
239
|
+
|
|
240
|
+
## Convergence And Stop Rules
|
|
241
|
+
|
|
242
|
+
Apply protocol repair, no-progress, stop, and waiver semantics. The following
|
|
243
|
+
audit is implementation-specific.
|
|
244
|
+
|
|
245
|
+
Before launching more reviewers, run one non-terminal convergence audit after
|
|
246
|
+
two consecutive Closure waves in one gate, ten total implementation review
|
|
247
|
+
passes, or 90 minutes when timing is available. One coordinated launch over all
|
|
248
|
+
affected lineages is one wave regardless of parallel pass count.
|
|
249
|
+
|
|
250
|
+
Persist one audit epoch with trigger, triggering pass/wave/revision,
|
|
251
|
+
completion, dispositions, selected sessions, and resume reason. It survives
|
|
252
|
+
resume. After a material repair/evidence change proves progress, rearm by
|
|
253
|
+
recording the current total pass count, current gate-local Closure-wave count,
|
|
254
|
+
and current timestamp as new baselines. The next audit opens only after a
|
|
255
|
+
post-rearm delta reaches two Closure waves in that gate, ten implementation
|
|
256
|
+
review passes, or 90 minutes; already-consumed counts or time cannot reopen it
|
|
257
|
+
immediately. Without progress do not rearm and use the existing stop rules.
|
|
258
|
+
Thresholds never approve, waive, downgrade, or block by themselves, and
|
|
259
|
+
distinct failure mechanics remain distinct even when they protect one
|
|
260
|
+
invariant.
|
|
261
|
+
|
|
262
|
+
Return `Approved` for the final settled revision only when protocol state is
|
|
263
|
+
`clear` and:
|
|
264
|
+
|
|
265
|
+
- every mandatory lens has independent coverage
|
|
266
|
+
- final validation and required cleanup/code-review gates ran
|
|
267
|
+
|
|
268
|
+
Map protocol `stopped` to `Blocked`. Implementation-specific blockers also
|
|
269
|
+
include:
|
|
270
|
+
|
|
271
|
+
- a defect reopens repeatedly and exposes a source-of-truth or repair-design
|
|
272
|
+
contradiction that root cannot resolve from current evidence
|
|
273
|
+
- an execution risk remains open without an explicit user decision to accept it
|
|
274
|
+
|
|
275
|
+
Do not mark implementation `Blocked` merely because review has run several
|
|
276
|
+
times. Resolve the repair, evidence, or decision problem first.
|
|
277
|
+
|
|
278
|
+
Map protocol `waived` to `Waived` and preserve skipped coverage and open risks.
|
|
279
|
+
It remains non-approval; target authority or downstream policy may still block
|
|
280
|
+
delivery.
|
|
281
|
+
|
|
282
|
+
## Required Handoff
|
|
283
|
+
|
|
284
|
+
Use the protocol result envelope and add:
|
|
285
|
+
|
|
286
|
+
```text
|
|
287
|
+
Implementation Review Profile: <simple | medium | high>
|
|
288
|
+
Review Outcome: <Approved | Blocked | Waived>
|
|
289
|
+
Authority Artifact: <approved spec path>
|
|
290
|
+
Implementation Checkpoint: <checkpoint or final>
|
|
291
|
+
```
|
|
292
|
+
|
|
293
|
+
## Contract Test Ledger
|
|
294
|
+
|
|
295
|
+
| Invariant | Risk It Prevents | First Test / Proof | Status |
|
|
296
|
+
| --- | --- | --- | --- |
|
|
297
|
+
| Review pass counts are audit metrics, while one durable Review Plan and Defect Ledger span the entire spec. | Each slice silently recreates a new loop or a repairable spec blocks on an arbitrary count. | Manual eval scenario 15 | planned |
|
|
298
|
+
| Intermediate checkpoints never trigger cleanup; final review is one settled profile-selected wave, and separate cleanup is exceptional. | Per-slice hygiene or size-driven cleanup adds latency and restarts review over unstable work. | Manual eval scenario 17 | planned |
|
|
299
|
+
| Audit epochs persist across resume and rearm only after material progress. | Thresholds repeatedly trigger audits or become terminal limits. | Manual eval scenario 12 | planned |
|
|
300
|
+
|
|
301
|
+
Keep these rows `planned` until the corresponding operator eval is run and its
|
|
302
|
+
result is saved.
|
|
@@ -0,0 +1,49 @@
|
|
|
1
|
+
# Review Gates
|
|
2
|
+
|
|
3
|
+
This file owns applicability only. Approved-spec execution follows
|
|
4
|
+
[`implementation-review-loop.md`](implementation-review-loop.md), which owns
|
|
5
|
+
checkpoint order, Full/Closure topology, durable state, and final coverage.
|
|
6
|
+
Direct work uses the gates below without manufacturing Module state.
|
|
7
|
+
|
|
8
|
+
## Cleanup Review
|
|
9
|
+
|
|
10
|
+
Do not launch a separate `$cleanup-review` from size or risk classification
|
|
11
|
+
alone. The final `$code-review` spec/standards lens owns bounded cleanup for
|
|
12
|
+
simple, medium, large, and high-risk changes; high runs that lens in its own
|
|
13
|
+
parallel reviewer track.
|
|
14
|
+
|
|
15
|
+
Use separate `$cleanup-review` only when the user, approved source, or repo
|
|
16
|
+
policy names a concrete evidenced simplification risk that cannot fit the
|
|
17
|
+
bounded spec/standards lens. It runs before final `$code-review`; approved-spec
|
|
18
|
+
execution lets the shared Module schedule it.
|
|
19
|
+
|
|
20
|
+
Treat a change as large only when it contains several independently verifiable
|
|
21
|
+
runtime workflows or material cross-owner, cross-repo, release-sequencing, or
|
|
22
|
+
rollback coordination. File count, module count, or a broad mechanical diff is
|
|
23
|
+
not enough. Default a coherent feature with one behavior and one validation
|
|
24
|
+
path to medium even when it touches several files.
|
|
25
|
+
|
|
26
|
+
Cleanup is not a correctness review. Its skill owns the detailed lens,
|
|
27
|
+
confidence handling, repair integration, and output contract.
|
|
28
|
+
|
|
29
|
+
## Final Code Review
|
|
30
|
+
|
|
31
|
+
Run final `$code-review` when implementation changes include any of:
|
|
32
|
+
|
|
33
|
+
- a medium/large feature or shared multi-module business behavior;
|
|
34
|
+
- API/DTO/schema, migration, persistence, auth, permission, payment, cache,
|
|
35
|
+
concurrency, background-job, or shared-state contracts;
|
|
36
|
+
- shared UI/navigation/middleware/core flows;
|
|
37
|
+
- runtime logic across three or more files when the change is behavioral rather
|
|
38
|
+
than mechanical and crosses one owner or validation seam.
|
|
39
|
+
|
|
40
|
+
Do not invoke it automatically for documentation/copy/comments, tests-only or
|
|
41
|
+
styling-only changes, formatting/renames, mechanical refactors, or isolated
|
|
42
|
+
one-file fixes with low regression risk.
|
|
43
|
+
|
|
44
|
+
`$code-review` owns reviewer topology, confidence, auto-fix, and output rules.
|
|
45
|
+
Its spec/standards reviewer owns the bounded cleanup lens for every profile; do
|
|
46
|
+
not launch `$cleanup-review` for the same settled diff without an explicit
|
|
47
|
+
concrete reason beyond size or risk labels.
|
|
48
|
+
For approved specs, required final coverage remains part of the shared
|
|
49
|
+
Implementation Review State.
|
|
@@ -0,0 +1,170 @@
|
|
|
1
|
+
# Review Protocol
|
|
2
|
+
|
|
3
|
+
This Module owns review mechanics shared by artifact and implementation review:
|
|
4
|
+
capsules, Full/Closure, reviewer lineage, the canonical Defect Ledger, no-progress,
|
|
5
|
+
stop/waiver semantics, and the result envelope. Target Modules own authority,
|
|
6
|
+
risk/profile selection, topology, durable orchestration, gate order, and final
|
|
7
|
+
outcome mapping. Callers select a target Module and never copy this protocol.
|
|
8
|
+
|
|
9
|
+
## Interface
|
|
10
|
+
|
|
11
|
+
```text
|
|
12
|
+
apply_review_protocol(
|
|
13
|
+
review_capsule,
|
|
14
|
+
reviewer_lineages
|
|
15
|
+
) -> protocol_result
|
|
16
|
+
```
|
|
17
|
+
|
|
18
|
+
`review_capsule` is the normalized Full/Closure capsule defined below. The result
|
|
19
|
+
contains `state: clear | repair-required | stopped | waived`, pass
|
|
20
|
+
mode/revision/lineage/session/lenses, mandatory coverage, ledger transitions, and any
|
|
21
|
+
stop or waiver record. The capsule's canonical Defect Ledger is the only ledger
|
|
22
|
+
input. `clear` requires assigned-scope coverage and every non-superseded blocker
|
|
23
|
+
or execution risk to be `verified`, except an explicitly `accepted-risk`
|
|
24
|
+
execution risk. A superseded record counts as resolved only when `superseded_by`
|
|
25
|
+
resolves through an acyclic chain of distinct same-ledger IDs to one
|
|
26
|
+
non-superseded canonical replacement; that terminal replacement controls
|
|
27
|
+
clearance. A missing, self-referential, or cyclic replacement chain prevents
|
|
28
|
+
`clear`. The target Module decides whether `clear` is enough for `Approved`. A
|
|
29
|
+
reviewer lineage binds one activation, profile/lenses, Full coverage, and defect
|
|
30
|
+
IDs; physical sessions may rotate without creating a new activation or Full.
|
|
31
|
+
|
|
32
|
+
## Review Capsule
|
|
33
|
+
|
|
34
|
+
Every fresh reviewer receives a bounded self-contained capsule:
|
|
35
|
+
|
|
36
|
+
- target kind/path, pinned revision, profile, mode, and assigned lenses
|
|
37
|
+
- one review question and why existing valid coverage does not answer it
|
|
38
|
+
- goal, authority, approved scope, and relevant source references
|
|
39
|
+
- required Evidence Index entries, compact validation, and verification gaps
|
|
40
|
+
- canonical Defect Ledger
|
|
41
|
+
- for Closure: repair diff, affected contracts, and Revision Map
|
|
42
|
+
`changed target -> defect IDs -> affected invariants/contracts`
|
|
43
|
+
|
|
44
|
+
Exclude raw parent history, old targets, prior reviewer prose, repeated logs,
|
|
45
|
+
unrelated output, and broad inventories. Reviewers may inspect extra evidence
|
|
46
|
+
needed for a finding. Reuse unchanged evidence; invalidate only entries affected
|
|
47
|
+
by changed targets/contracts or stale external sources. External evidence keeps
|
|
48
|
+
its direct source and retrieval date.
|
|
49
|
+
|
|
50
|
+
Before launch, reuse clear coverage for the same target revision, question, and
|
|
51
|
+
lenses. Do not launch a reviewer whose question is already answered; a changed
|
|
52
|
+
target or a distinct artifact-versus-implementation question is new coverage.
|
|
53
|
+
|
|
54
|
+
## Review Modes
|
|
55
|
+
|
|
56
|
+
**Full:** read the complete pinned target required by assigned lenses and return
|
|
57
|
+
all visible evidence-backed blockers/execution risks as one batch.
|
|
58
|
+
|
|
59
|
+
**Closure:** use only affected reviewer lineages to verify repaired IDs, the
|
|
60
|
+
repair diff, and contract fan-out. Do not re-audit unchanged areas. The first
|
|
61
|
+
Closure normally reuses the Full session. If it reopens a defect and another
|
|
62
|
+
repair follows, run the next Closure in a fresh session of the same lineage.
|
|
63
|
+
Before a Closure launch, rotate earlier after compaction/interruption or when
|
|
64
|
+
observable context usage reaches 40%. The fresh session receives the bounded
|
|
65
|
+
capsule, remains Closure, and preserves Full coverage and defect authority. One
|
|
66
|
+
coordinated launch over affected lineages is one Closure wave.
|
|
67
|
+
|
|
68
|
+
Closure admits a new defect only when it names the repair target/contract that
|
|
69
|
+
introduced or materially widened the trigger, or a high-confidence
|
|
70
|
+
`blocker`/`execution-risk` violates a source-required invariant in the assigned
|
|
71
|
+
lens. Each newly admitted blocker or execution risk starts as `status: open`
|
|
72
|
+
and `disposition: repair-now`, then moves `open -> fixed -> verified` only
|
|
73
|
+
through repair and independent affected-lens verification. Pre-existing
|
|
74
|
+
adjacent issues are improvements. Verification stays with an affected session
|
|
75
|
+
or an already planned covering Full reviewer. Start a new Full only when the
|
|
76
|
+
repair invalidated mandatory-lens coverage.
|
|
77
|
+
|
|
78
|
+
## Canonical Defect Ledger
|
|
79
|
+
|
|
80
|
+
Root assigns stable IDs and deduplicates only when invariant and failure
|
|
81
|
+
mechanics both match. Wording changes do not create defects.
|
|
82
|
+
|
|
83
|
+
```yaml
|
|
84
|
+
id: REVIEW-CONC-003
|
|
85
|
+
class: blocker | execution-risk | improvement
|
|
86
|
+
disposition: repair-now | planned-final-verification | follow-up-improvement | accepted-risk
|
|
87
|
+
status: open | fixed | verified | blocked | reopened | accepted-risk | superseded
|
|
88
|
+
severity: critical | high | medium | low
|
|
89
|
+
confidence: high | medium | low
|
|
90
|
+
invariant: "<observable rule>"
|
|
91
|
+
failure: "<concrete failure mechanics>"
|
|
92
|
+
evidence: ["<target/repo/source reference>"]
|
|
93
|
+
repair: "<smallest sufficient change>"
|
|
94
|
+
affected_targets: ["<path, section, contract, or lens>"]
|
|
95
|
+
introduced_in_review: 2
|
|
96
|
+
transition_history: []
|
|
97
|
+
reopen_count: 0
|
|
98
|
+
acceptance: null | { authority, reason, scope, target_revision }
|
|
99
|
+
superseded_by: null | "<canonical replacement defect ID>"
|
|
100
|
+
fixed_in_revision: null
|
|
101
|
+
verified_in_review: null
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
Reviewers reuse supplied IDs; new candidates use `NEW-<LENS>-NN` until root
|
|
105
|
+
deduplicates them. Root may mark a duplicate `superseded` only after recording
|
|
106
|
+
its canonical replacement in `superseded_by`. The chain must be acyclic and end
|
|
107
|
+
at a distinct non-superseded record; a superseded record never hides the
|
|
108
|
+
terminal replacement's lifecycle. Disposition and lifecycle are separate.
|
|
109
|
+
`fixed` requires a later affected-lens Closure or planned Full to become
|
|
110
|
+
`verified`; tests alone do not independently verify review findings.
|
|
111
|
+
|
|
112
|
+
Improvements use `follow-up-improvement` and do not block. Only an execution
|
|
113
|
+
risk may become `accepted-risk`, after explicit authority, reason, scope, and
|
|
114
|
+
target revision are recorded. A blocker cannot be accepted or downgraded.
|
|
115
|
+
Target Modules decide where `planned-final-verification` is legal; it remains
|
|
116
|
+
open until the scheduled lens verifies it.
|
|
117
|
+
|
|
118
|
+
## Repair, Stop, And Waiver
|
|
119
|
+
|
|
120
|
+
Root repairs compatible findings in one batch. Before another Closure, target
|
|
121
|
+
revision/strategy, relevant evidence, defect state, or source decision must
|
|
122
|
+
change materially. Otherwise root must change the repair, prove the finding
|
|
123
|
+
invalid, or surface the decision preventing convergence.
|
|
124
|
+
|
|
125
|
+
Pass counts and elapsed time are audit signals, never approval, waiver,
|
|
126
|
+
downgrade, or blocking conditions. Target Modules may define audit epochs while
|
|
127
|
+
preserving this rule.
|
|
128
|
+
|
|
129
|
+
A bounded poll timeout while the reviewer session remains live is non-terminal:
|
|
130
|
+
report that the reviewer has not completed within the polling window, not that
|
|
131
|
+
it is stuck, failed, or unavailable. Root may send at most one conclude request
|
|
132
|
+
for that reviewer turn. Further empty polls neither authorize another conclude
|
|
133
|
+
request nor replacement/cancellation. Replace or cancel only after the reviewer
|
|
134
|
+
session or transport explicitly reports `failed`, `lost`, or unavailable; a
|
|
135
|
+
late usable result from the original live session remains authoritative. Persist
|
|
136
|
+
session `live|timed-out` state and `conclude_requested_at` so resume preserves
|
|
137
|
+
the limit.
|
|
138
|
+
|
|
139
|
+
Return `stopped` without another review when repair needs a product/scope/owner
|
|
140
|
+
decision, mandatory evidence or reviewer is unavailable, no substantive repair
|
|
141
|
+
exists, or the same defect repeats without progress.
|
|
142
|
+
|
|
143
|
+
Return `waived` only after explicit user instruction. Record skipped coverage
|
|
144
|
+
and open defects. Waiver never means approval, verifies/accepts no defect, and
|
|
145
|
+
skips no separate downstream gate. The target Module maps it to `Waived` or
|
|
146
|
+
`Blocked` using its authority rules.
|
|
147
|
+
|
|
148
|
+
## Result Envelope
|
|
149
|
+
|
|
150
|
+
```text
|
|
151
|
+
Review Passes: <total; Full; Closure; fresh sessions when tracked>
|
|
152
|
+
Mandatory Coverage: <covered lenses or gaps>
|
|
153
|
+
Verified Defects: <IDs or None>
|
|
154
|
+
Accepted Risks: <IDs, authority, and reason or None>
|
|
155
|
+
Open Defects: <IDs or None>
|
|
156
|
+
```
|
|
157
|
+
|
|
158
|
+
Target Modules add profile, mapped outcome, authority, Adapter verdict,
|
|
159
|
+
checkpoint, status, or performance fields without redefining common fields.
|
|
160
|
+
|
|
161
|
+
## Contract Test Ledger
|
|
162
|
+
|
|
163
|
+
| Invariant | Risk It Prevents | First Test / Proof | Status |
|
|
164
|
+
| --- | --- | --- | --- |
|
|
165
|
+
| Full/Closure stay causal and affected-lens-only within reviewer lineage | Broad review restarts, stale context, or lost defect authority | Evals 11, 16 | green |
|
|
166
|
+
| Every launch answers an uncovered question on its pinned target | Duplicate reviewers repeat valid coverage without new evidence | Eval 23 | green |
|
|
167
|
+
| Both consumers use one defect schema and lifecycle | Artifact and implementation defects drift | Eval 12 | green |
|
|
168
|
+
| No-progress requires material change | Unchanged Closure repeats or numeric limits become terminal | Eval 12 | green |
|
|
169
|
+
| A live reviewer poll timeout stays non-terminal and permits at most one conclude request | Root falsely reports a hang, cancels useful work, or launches a duplicate reviewer | Eval 12 | green |
|
|
170
|
+
| Waiver skips coverage without accepting defects | Skipped review is reported as approval or risk acceptance | Eval 12 | green |
|