okstra 0.171.0 → 0.173.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (106) hide show
  1. package/README.md +8 -6
  2. package/docs/architecture/storage-model.md +11 -0
  3. package/docs/architecture.md +29 -14
  4. package/docs/cli.md +40 -7
  5. package/docs/for-ai/skills/okstra-user-response.md +2 -2
  6. package/docs/performance-improvement-plan-v2.md +6 -5
  7. package/docs/project-structure-overview.md +24 -14
  8. package/docs/task-process/README.md +5 -3
  9. package/docs/task-process/error-analysis.md +2 -2
  10. package/docs/task-process/final-verification.md +2 -2
  11. package/docs/task-process/implementation-option-selection.md +70 -0
  12. package/docs/task-process/implementation-planning.md +23 -15
  13. package/docs/task-process/requirements-discovery.md +2 -2
  14. package/package.json +1 -1
  15. package/runtime/BUILD.json +2 -2
  16. package/runtime/agents/workers/report-writer-worker.md +30 -6
  17. package/runtime/bin/lib/okstra/cli.sh +5 -1
  18. package/runtime/bin/lib/okstra/globals.sh +1 -0
  19. package/runtime/bin/lib/okstra/usage.sh +3 -0
  20. package/runtime/bin/okstra.sh +2 -0
  21. package/runtime/prompts/duties/direction-selection-worker.md +44 -0
  22. package/runtime/prompts/duties/planning-worker.md +12 -4
  23. package/runtime/prompts/launch.template.md +4 -0
  24. package/runtime/prompts/lead/adapters/cmux.md +1 -1
  25. package/runtime/prompts/lead/context-loader.md +1 -1
  26. package/runtime/prompts/lead/convergence.md +5 -5
  27. package/runtime/prompts/lead/okstra-lead-contract.md +42 -17
  28. package/runtime/prompts/lead/plan-body-verification.md +42 -14
  29. package/runtime/prompts/lead/report-writer.md +38 -15
  30. package/runtime/prompts/lead/team-contract.md +2 -0
  31. package/runtime/prompts/profiles/_clarification-recommendation.md +3 -1
  32. package/runtime/prompts/profiles/_common-contract.md +3 -2
  33. package/runtime/prompts/profiles/_implementation-deliverable.md +2 -2
  34. package/runtime/prompts/profiles/error-analysis.md +3 -3
  35. package/runtime/prompts/profiles/final-verification.md +3 -3
  36. package/runtime/prompts/profiles/forbidden-actions.json +7 -0
  37. package/runtime/prompts/profiles/implementation-option-selection.md +35 -0
  38. package/runtime/prompts/profiles/implementation-planning.md +56 -37
  39. package/runtime/prompts/profiles/implementation.md +2 -1
  40. package/runtime/prompts/profiles/improvement-discovery.md +1 -1
  41. package/runtime/prompts/profiles/requirements-discovery.md +3 -3
  42. package/runtime/prompts/wizard/prompts.ko.json +9 -1
  43. package/runtime/python/okstra_ctl/adapters/hosts/antigravity/relay.md +1 -1
  44. package/runtime/python/okstra_ctl/adapters/hosts/claude-code/relay.md +1 -1
  45. package/runtime/python/okstra_ctl/adapters/hosts/codex/relay.md +1 -1
  46. package/runtime/python/okstra_ctl/adapters/hosts/external/relay.md +1 -1
  47. package/runtime/python/okstra_ctl/adapters/hosts/grok/relay.md +1 -1
  48. package/runtime/python/okstra_ctl/adapters/hosts/kimi/relay.md +1 -1
  49. package/runtime/python/okstra_ctl/agent_activity.py +306 -0
  50. package/runtime/python/okstra_ctl/agent_invocation.py +1 -0
  51. package/runtime/python/okstra_ctl/analysis_packet.py +6 -0
  52. package/runtime/python/okstra_ctl/clarification_items.py +37 -20
  53. package/runtime/python/okstra_ctl/exact_coverage.py +128 -0
  54. package/runtime/python/okstra_ctl/fix_cycles.py +3 -1
  55. package/runtime/python/okstra_ctl/implementation_direction.py +836 -0
  56. package/runtime/python/okstra_ctl/implementation_options.py +479 -0
  57. package/runtime/python/okstra_ctl/lead_events.py +47 -4
  58. package/runtime/python/okstra_ctl/plan_items.py +51 -3
  59. package/runtime/python/okstra_ctl/render.py +12 -3
  60. package/runtime/python/okstra_ctl/render_final_report.py +1 -0
  61. package/runtime/python/okstra_ctl/report_contract.py +45 -13
  62. package/runtime/python/okstra_ctl/report_finalize.py +51 -14
  63. package/runtime/python/okstra_ctl/report_html/common.py +5 -3
  64. package/runtime/python/okstra_ctl/report_html/render.py +4 -2
  65. package/runtime/python/okstra_ctl/report_html/router.py +4 -0
  66. package/runtime/python/okstra_ctl/report_html/view_models/implementation_option_selection.py +32 -0
  67. package/runtime/python/okstra_ctl/report_html/view_models/implementation_planning.py +42 -11
  68. package/runtime/python/okstra_ctl/report_translation.py +14 -0
  69. package/runtime/python/okstra_ctl/report_views.py +148 -12
  70. package/runtime/python/okstra_ctl/run.py +350 -2
  71. package/runtime/python/okstra_ctl/scope_provenance.py +15 -9
  72. package/runtime/python/okstra_ctl/user_response.py +75 -0
  73. package/runtime/python/okstra_ctl/wizard.py +144 -0
  74. package/runtime/python/okstra_ctl/worker_audit_ledger.py +150 -0
  75. package/runtime/python/okstra_ctl/worker_prompt_policy.py +2 -0
  76. package/runtime/python/okstra_ctl/workflow.py +29 -7
  77. package/runtime/schemas/final-report-v2.0.schema.json +1623 -143
  78. package/runtime/skills/okstra-user-response/SKILL.md +2 -2
  79. package/runtime/templates/reports/final-report-v2.template.md +12 -0
  80. package/runtime/templates/reports/final-verification-input.template.md +1 -1
  81. package/runtime/templates/reports/html/assets/base.css +7 -0
  82. package/runtime/templates/reports/html/base.template.html +3 -2
  83. package/runtime/templates/reports/html/i18n/en.json +27 -2
  84. package/runtime/templates/reports/html/i18n/ko.json +27 -2
  85. package/runtime/templates/reports/html/macros/forms.html +42 -4
  86. package/runtime/templates/reports/html/tasks/implementation-option-selection.template.html +49 -0
  87. package/runtime/templates/reports/html/tasks/implementation-planning.template.html +61 -2
  88. package/runtime/templates/reports/i18n/en.json +17 -0
  89. package/runtime/templates/reports/implementation-input.template.md +4 -2
  90. package/runtime/templates/reports/implementation-planning-input.template.md +18 -4
  91. package/runtime/templates/reports/improvement-discovery-input.template.md +1 -1
  92. package/runtime/templates/reports/md/tasks/implementation-option-selection.template.md +13 -0
  93. package/runtime/templates/reports/md/tasks/implementation-planning.template.md +17 -0
  94. package/runtime/templates/reports/report.js +137 -21
  95. package/runtime/templates/reports/task-brief.template.md +9 -3
  96. package/runtime/templates/reports/user-response.template.md +28 -5
  97. package/runtime/templates/worker-prompt-preamble.md +16 -0
  98. package/runtime/validators/validate-implementation-plan-stages.py +106 -1
  99. package/runtime/validators/validate-report-views.py +2 -2
  100. package/runtime/validators/validate-run.py +1124 -54
  101. package/runtime/validators/validate_improvement_report.py +5 -1
  102. package/runtime/validators/validate_session_conformance.py +523 -35
  103. package/src/cli-registry.mjs +7 -0
  104. package/src/commands/execute/codex-run.mjs +1 -0
  105. package/src/commands/execute/render-bundle.mjs +1 -0
  106. package/src/commands/report/agent-activity.mjs +21 -0
@@ -0,0 +1,44 @@
1
+ ---
2
+ id: direction-selection-worker
3
+ version: 1
4
+ kind: role
5
+ appliesTo: direction-selection-worker
6
+ ---
7
+
8
+ # Direction Selection Worker Duty Contract
9
+
10
+ ## Responsibility
11
+
12
+ Compare feasible directions before planning in `candidate-comparison` mode, or validate one preselected direction in `preselected-validation` mode.
13
+
14
+ ## Required conduct
15
+
16
+ In `candidate-comparison` mode, inspect the evidence needed to distinguish candidates, submit no more than three candidates, state the strongest counterevidence for each one, and map every candidate to the stable brief end-state IDs it satisfies, preserves, or leaves unresolved. In `preselected-validation` mode, validate the one preselected direction against that evidence and mapping; the worker must not generate new candidates.
17
+
18
+ ## Decision principles
19
+
20
+ Score candidates against the same stated criteria. Prefer evidence-backed feasibility over familiarity, and preserve a rejected candidate when its evidence or trade-off could affect the later planning decision.
21
+
22
+ ## Authority and boundaries
23
+
24
+ Select directions only. Do not edit project state, author detailed file lists, create stage maps, prescribe execution commands, or approve an implementation plan.
25
+
26
+ ## Evidence standard
27
+
28
+ Each candidate, score, counterexample, and requirement mapping cites inspected evidence. State uncertainty when the code or brief cannot establish a criterion.
29
+
30
+ ## Collaboration contract
31
+
32
+ Reason independently from other workers. Do not collapse overlapping candidates or seek agreement before convergence; provide the evidence that lets the lead merge and re-evaluate them.
33
+
34
+ ## Completion criteria
35
+
36
+ In `candidate-comparison` mode, every submitted candidate has a criterion score, counterevidence, and stable requirement mapping. In `preselected-validation` mode, the one preselected direction has a validation result, counterevidence, and stable requirement mapping. Rejected candidates retain their audit reason and evidence.
37
+
38
+ ## Forbidden conduct
39
+
40
+ Do not turn a candidate into a detailed implementation plan, invent a requirement ID, omit contrary evidence, present a selection as user approval, or generate a new candidate in `preselected-validation` mode.
41
+
42
+ ## Blocked-state reporting
43
+
44
+ Name the missing evidence or unresolved requirement that prevents a candidate comparison, the inspection attempted, and the criterion it leaves unscored.
@@ -9,11 +9,19 @@ appliesTo: planning-worker
9
9
 
10
10
  ## Responsibility
11
11
 
12
- Produce an implementation direction a person can approve: feasible options with their trade-offs, one recommendation, and stages that carry the requirement to a verifiable end — all without writing the implementation.
12
+ Produce an executable implementation plan without writing the implementation.
13
+
14
+ ### Selected-direction responsibility
15
+
16
+ When `selected-direction.json` is present, read it and the requirements ledger first. Preserve the selected direction's core mechanism, architecture boundaries, user constraints, and planning invariants. Concretize only its files, interfaces, stages, validation, and rollback. Link every file and stage back to original requirements, and link every original requirement forward to its files, stages, and checks. If current code evidence requires changing the direction, return `direction-invalidated` with evidence and stop. Direction selection remains upstream of this branch.
17
+
18
+ ### Legacy candidate-comparison compatibility
19
+
20
+ For a legacy rerun without `selected-direction.json`, retain Option Candidates, trade-offs, the Recommended Option, and `P-Opt-*` verification semantics. Only this compatibility branch compares alternatives or chooses a recommendation.
13
21
 
14
22
  ## Required conduct
15
23
 
16
- Read the current state of the code the work touches before drafting options; compare at least two feasible options on evidence from that code unless the decision is already settled upstream; tie the recommendation to the trade-off that decides it; split the work into stages along real dependencies with each stage's validation signal and rollback; and connect every requirement to the stage that satisfies it.
24
+ Read the current state of the code the work touches before planning. Split work into stages along real dependencies with each stage's validation signal and rollback. Connect every requirement to the stage that satisfies it. In the selected-direction branch, verify direction preservation before adding plan detail. In the legacy branch, compare options on current-code evidence.
17
25
 
18
26
  ## Decision principles
19
27
 
@@ -29,11 +37,11 @@ Every cited path, symbol, and command must exist as written and be executable in
29
37
 
30
38
  ## Collaboration contract
31
39
 
32
- Draft independently of the other planners rather than converging on the first option proposed. Leave the choice between competing plans and the resolution of contested items to convergence and the lead, and hand the executor a plan complete enough to follow without re-deriving the decisions behind it.
40
+ Draft independently of the other planners. Leave contested plan details to convergence and the lead, and hand the executor a plan complete enough to follow without re-deriving the selected direction.
33
41
 
34
42
  ## Completion criteria
35
43
 
36
- Options, trade-offs, the recommendation, the stages with their dependencies, validation and rollback, and requirement coverage are all present and mutually consistent; every unresolved decision is recorded as such rather than assumed; and no stage depends on work the plan never places.
44
+ For selected-direction planning, the snapshot and requirements ledger are preserved; files, interfaces, stages, validation, rollback, and requirements links are mutually consistent; planning invariants have evidence; and no hidden direction change is present. For legacy candidate-comparison compatibility, Option Candidates, trade-offs, the Recommended Option, and `P-Opt-*` semantics remain present. Every unresolved decision is recorded rather than assumed, and no stage depends on work the plan never places.
37
45
 
38
46
  ## Forbidden conduct
39
47
 
@@ -7,6 +7,10 @@
7
7
 
8
8
  Emit one `PROGRESS: <phase-id> <verb-phrase>` line as plain user-facing text at every checkpoint enumerated in the lifecycle core contract (`{{OKSTRA_LEAD_CONTRACT_PATH}}` "Progress reporting (BLOCKING)") — phase-1-intake start/complete, phase-2-prompts, phase-3-team-create, phase-4-dispatch (per worker), phase-5-collect (per worker), phase-5.5-convergence (per round), phase-6-synthesis, phase-7-persist, and final `complete`. One line per checkpoint, never batched, never replaced with prose. This is the only signal the user has during multi-minute silent windows.
9
9
 
10
+ When the run manifest declares `activityContractVersion: 1`, call `okstra agent-activity append` before each required activity boundary. Only after the structured append succeeds, emit the matching `PROGRESS:` line and the immediately following `ACTIVITY:` projection from the same fields. If the structured append fails, do not mark that boundary completed. Never reconstruct structured activity by parsing `ACTIVITY:` conversation text.
11
+
12
+ For a new `implementation-planning` run, the plan-body sequence is initial verification → one planner self-fix → targeted re-verification → user gate. The initial verification is round 1, the targeted re-verification is round 2, and a second automatic self-fix is a contract violation. A user-directed correction does not consume the automatic self-fix limit, and a verification failure after that correction does not restart the automatic loop.
13
+
10
14
  ## Current Phase Boundary
11
15
 
12
16
  - Current lifecycle phase: `{{WORKFLOW_CURRENT_PHASE}}`
@@ -31,7 +31,7 @@ It overrides only the worker-dispatch portion of the selected host relay, not th
31
31
  | `await_workers` | Run `okstra team await --project-root <root> --run-manifest <path>` through the host's asynchronous shell facility. |
32
32
  | `redispatch_worker` | Create the core-specified fresh jobs file and dispatch it with a new `dispatchKind`; never reuse a live worker conversation. |
33
33
  | `shutdown_workers` | Run `okstra team teardown --project-root <root> --run-manifest <path>` only after the user-approved cleanup gate. |
34
- | `record_lead_event` | Append the required structured event to the manifest-provided `leadEventsPath`; emit the matching user-facing `PROGRESS:` line. |
34
+ | `record_lead_event` | Append progress and activity records to the manifest-provided `leadEventsPath`. Emit the matching `PROGRESS:` line and, when an activity record is required, the immediately following `ACTIVITY:` line from the same structured fields. |
35
35
  | `collect_usage` | Collect artifact/CLI-log-backed usage through the existing Okstra token-usage path; never substitute another runtime's session log. |
36
36
 
37
37
  ## Pane placement is not yours to compute
@@ -37,7 +37,7 @@
37
37
  | `projectId` | Project ID |
38
38
  | `taskGroup` | Task group |
39
39
  | `taskId` | Task ID |
40
- | `taskType` | Analysis type (requirements-discovery, error-analysis, implementation-planning, implementation, final-verification, release-handoff, plus the sidetrack improvement-discovery) |
40
+ | `taskType` | Analysis type (requirements-discovery, error-analysis, implementation-option-selection, implementation-planning, implementation, final-verification, release-handoff, plus the sidetrack improvement-discovery) |
41
41
  | `workCategory` | bugfix / feature / improvement / refactor / ops / unknown |
42
42
  | `recommendedWorkers` | List of selected workers |
43
43
  | `currentStatus` | Current task status |
@@ -22,7 +22,7 @@
22
22
 
23
23
  ## Scope and Terminology (BLOCKING)
24
24
 
25
- This contract governs **Phase 5.5 (Convergence loop)** — a *lead operating phase* inside a single okstra run, not a task-type lifecycle phase. It leaves the 6 task-type lifecycle phases (`requirements-discovery` → `error-analysis` → `implementation-planning` → `implementation` → `final-verification` → `release-handoff`, see [okstra-lead-contract](./okstra-lead-contract.md) "Lifecycle Phase Boundaries") unchanged; the lead operating phases (Phase 1 Intake → Phase 7 Persist, see [okstra-lead-contract](./okstra-lead-contract.md) "Quick Reference") drive a *single* task-type run.
25
+ This contract governs **Phase 5.5 (Convergence loop)** — a *lead operating phase* inside a single okstra run, not a task-type lifecycle phase. It leaves the 7 task-type lifecycle phases (`requirements-discovery` → `error-analysis` → `implementation-option-selection` → `implementation-planning` → `implementation` → `final-verification` → `release-handoff`, see [okstra-lead-contract](./okstra-lead-contract.md) "Lifecycle Phase Boundaries") unchanged; the lead operating phases (Phase 1 Intake → Phase 7 Persist, see [okstra-lead-contract](./okstra-lead-contract.md) "Quick Reference") drive a *single* task-type run.
26
26
 
27
27
  **`contested` is a final classification only.** It is NEVER an intermediate queue label. The verification queue carries findings that are *unique to a single worker* (entered in Round 0) or *mixed/unresolved after a re-verification round* (carried forward). The `contested` label is assigned only when the **last executed round** completes and the queue is still non-empty.
28
28
 
@@ -48,7 +48,7 @@ Configure this in the `convergence` block of `task-manifest.json`. If the block
48
48
  | `enabled` | `true` | If `false`, skip the convergence loop and use the existing consensus/divergence method |
49
49
  | `maxRounds` | phase-aware: `1` for `requirements-discovery`, `2` otherwise (range 1–3) | Maximum number of re-verification rounds. Discovery's routing/missing-input outputs gain little from a second round; other phases (especially `error-analysis`) keep `2`. Lead resolves the effective value when the manifest omits the key and records it in `config.effectiveMaxRounds` of the convergence state artifact. |
50
50
  | `verificationMode` | `"lightweight"` | `"lightweight"` or `"full-reanalysis"` |
51
- | `adversarial` | phase-aware: `true` for `requirements-discovery` / `error-analysis` / `implementation-planning` / `project-analysis` / `feature-analysis` / `change-impact-analysis`, `false` otherwise | When `true`, Phase 5.5 runs in **adversarial mode** (see §"Adversarial Verification Mode"): verifiers actively try to refute each finding, the burden of proof sits on the claim, and `verificationMode` is forced to `"full-reanalysis"` scoped to the finding's cited evidence. Resolved by `scripts/okstra_ctl/render.py` `_build_convergence_block` and recorded in `config.adversarial` of the convergence state artifact. |
51
+ | `adversarial` | phase-aware: `true` for `requirements-discovery` / `error-analysis` / `implementation-option-selection` / `implementation-planning` / `project-analysis` / `feature-analysis` / `change-impact-analysis`, `false` otherwise | When `true`, Phase 5.5 runs in **adversarial mode** (see §"Adversarial Verification Mode"): verifiers actively try to refute each finding, the burden of proof sits on the claim, and `verificationMode` is forced to `"full-reanalysis"` scoped to the finding's cited evidence. Resolved by `scripts/okstra_ctl/render.py` `_build_convergence_block` and recorded in `config.adversarial` of the convergence state artifact. |
52
52
 
53
53
  **Auto-disable rule (BLOCKING).** Convergence requires ≥2 analyser workers to produce a meaningful consensus tally. When the active profile's `Required workers:` block (see `prompts/profiles/*.md`) resolves to fewer than 2 analyser workers — e.g. `release-handoff` (zero analyser workers, lead-only) — the lead MUST treat `convergence.enabled` as `false` for that run regardless of manifest configuration, skip Phases 5.5 and the plan-body verification round ([plan-body-verification](./plan-body-verification.md)), and record `finalState: "converged"` with `totalRounds: 0`, `round2SkippedReason: "auto-disabled"`, an empty `roundHistory`, and an explanatory note in `config` (e.g. `"autoDisabled": "fewer-than-two-analysers"`). The plan-body round inherits the same rule via its `gating=false` advisory path.
54
54
 
@@ -160,7 +160,7 @@ Use each finding as a guide but reanalyze the original code/data yourself. High
160
160
 
161
161
  ## Adversarial Verification Mode
162
162
 
163
- Active only when `config.adversarial == true` (default for `requirements-discovery`, `error-analysis`, `implementation-planning`, `project-analysis`, `feature-analysis`, and `change-impact-analysis`; see §"Configuration"); when `false`, every rule in this section is inert and the collaborative behaviour elsewhere in this contract applies unchanged. In adversarial mode the verifier's job inverts: instead of confirming a peer's finding, the verifier **tries to break it**, and the burden of proof sits on the claim — a finding survives only if refutation attempts fail.
163
+ Active only when `config.adversarial == true` (default for `requirements-discovery`, `error-analysis`, `implementation-option-selection`, `implementation-planning`, `project-analysis`, `feature-analysis`, and `change-impact-analysis`; see §"Configuration"); when `false`, every rule in this section is inert and the collaborative behaviour elsewhere in this contract applies unchanged. In adversarial mode the verifier's job inverts: instead of confirming a peer's finding, the verifier **tries to break it**, and the burden of proof sits on the claim — a finding survives only if refutation attempts fail.
164
164
 
165
165
  ### Read-only analysis task contract
166
166
 
@@ -361,7 +361,7 @@ Lightweight reverify does not require the original `analysis-packet.md`, `analys
361
361
  - **Lightweight mode**: the clause directly contradicts the "Do NOT re-analyze the original source materials" instruction below. Including it forces workers to re-read the entire instruction-set per round per worker (3 workers × 2 rounds × 5+ files in the worst case) for no quality gain.
362
362
  - **Full-reanalysis mode**: workers DO need to re-read source materials, but only the analysis-worker file list (no `final-report-template.md`). If lead chooses to inject a reading clause here, it MUST mirror the audience-scoped enumeration in [okstra-lead-contract](./okstra-lead-contract.md) Phase 2 (no template).
363
363
 
364
- This is the single largest avoidable cost in `requirements-discovery`, `error-analysis`, and `implementation-planning` runs. Treat as mandatory.
364
+ This is the single largest avoidable cost in `requirements-discovery`, `error-analysis`, `implementation-option-selection`, and `implementation-planning` runs. Treat as mandatory.
365
365
 
366
366
  ### Lightweight Re-verification Prompt
367
367
 
@@ -565,7 +565,7 @@ Save it to `runs/<task-type>/state/convergence-<task-type>-<seq>.json`.
565
565
  Schema rules:
566
566
 
567
567
  - `schemaVersion`: literal string `"1.3"` for all new runs — both adversarial and collaborative. Historical readers accept `"1.0"` / `"1.1"` / `"1.2"` unchanged and never rewrite those artifacts during validation. v1.3 adds the strict coverage-critic ledger and rejects unknown top-level fields; work-state remains v1.0.
568
- - `config.adversarial`: boolean. `true` when this run used adversarial verification (default for `requirements-discovery` / `error-analysis` / `implementation-planning` / `project-analysis` / `feature-analysis` / `change-impact-analysis`). When `true`, `config.verificationMode` is `"full-reanalysis"` (scoped) and every `disagree` vote carries a non-null `disagreeBasis`.
568
+ - `config.adversarial`: boolean. `true` when this run used adversarial verification (default for `requirements-discovery` / `error-analysis` / `implementation-option-selection` / `implementation-planning` / `project-analysis` / `feature-analysis` / `change-impact-analysis`). When `true`, `config.verificationMode` is `"full-reanalysis"` (scoped) and every `disagree` vote carries a non-null `disagreeBasis`.
569
569
  - `config.effectiveMaxRounds`: the integer the lead actually used after resolving the phase-aware default (`1` for `requirements-discovery`, `2` otherwise). MUST equal `config.maxRounds` when the manifest explicitly set it.
570
570
  - `findings[].ticketIds`: array of ticket keys from Phase 4 grouping (parsed per the Round 0 step 5 rule). It is empty when the phase does not require ticket tagging; `"unknown"` is not a ticket key and must not be synthesized.
571
571
  - `findings[].rounds[].votes.<worker>.verdict`: enum, one of `agree | disagree | supplement | verification-error`. Lower-case tokens; map upper-case AGREE/DISAGREE/SUPPLEMENT verdicts emitted by workers to their lower-case form and map the input alias `unverifiable` to persisted `verification-error`. The latter represents either a terminal non-result dispatch or a completed dispatch that could not verify a particular finding (§"Worker failure handling in reverify"). Every vote has a non-empty `explanation`.
@@ -39,7 +39,7 @@ Read-side inspection (`/okstra-inspect`) and scheduling (`/okstra-schedule-gen`)
39
39
  | 5. Completion wait | Call `await_workers` and verify terminal state plus required artifacts | selected runtime adapter + `team-contract` |
40
40
  | 5.5 Convergence | Semantically group findings, then drive deterministic state transitions through `ConvergenceEngine` via `okstra convergence` | `convergence` |
41
41
  | 5.6 Critic pass | (opt-in) fresh one-shot critic pass through `redispatch_worker`: coverage gaps (discovery/error-analysis/impl-planning) or acceptance devil's-advocate (final-verification). The critic dispatch fires concurrently with the first 5.5 reverify round (its input is fixed at Round 0); gap/blocker verification (one round) completes here | `convergence` "Coverage critic pass" / "Acceptance critic pass" |
42
- | 6. Synthesis | Dispatch Report writer worker, review draft. **For `implementation-planning`: then run the Phase 6 plan-body verification sub-step (see Phase 6 section below).** | `report-writer` + `plan-body-verification` (sub-step) |
42
+ | 6. Synthesis | Dispatch Report writer worker, review draft. **For `implementation-planning`: then run the Phase 6 plan-body verification sub-step (see Phase 6 section below). Selected-direction plans verify `P-Dir-1`; legacy plans retain `P-Opt-*`.** | `report-writer` + `plan-body-verification` (sub-step) |
43
43
  | 7. Persist | Call `collect_usage`, update manifests, run the cleanup approval gate, then call `shutdown_workers` only on approval | selected runtime adapter + `report-writer` + this contract |
44
44
 
45
45
  ## Core operating contract
@@ -60,6 +60,7 @@ A single okstra run executes **exactly one** lifecycle phase. The phase is given
60
60
  |-----------------|-----------------|-------------------|
61
61
  | `requirements-discovery` | classification, routing decision, missing-input list, next-phase recommendation | code edits, plan documents, build/test execution that mutates state |
62
62
  | `error-analysis` | evidence, root-cause hypotheses, reproduction gaps, validation paths | code edits, implementation design, build/migration/deploy execution |
63
+ | `implementation-option-selection` | candidate comparison, counterevidence, criterion scores, requirement mappings, rejected-candidate audit | code edits, tests/builds, detailed file lists, stage maps, execution commands, plan approval |
63
64
  | `implementation-planning` | option matrix, trade-offs, dependencies, recommended order, validation/rollback strategy, Tier3 conformance scripts + manifest under the task-root `qa/` tree, **explicit user-approval request** | source code edits, file writes outside the run's `reports/`, `prompts/`, `state/`, `manifests/`, `worker-results/`, `status/`, `sessions/` directories and the task-root `qa/` tree, build/migration/deploy execution |
64
65
  | `implementation` | code edits authorised by an approved plan, accompanying tests | starting work without an approved `implementation-planning` final report carried in via `--clarification-response` or referenced in the brief |
65
66
  | `final-verification` | acceptance verdict, residual risk, regression notes; read-only execution of existing test/validation commands, run-artifact writes (qa result sidecars, `okstra handoff record-verified` on acceptance), and qaEnv-replica-only conformance runs are permitted | source code edits, refactors, scope expansion, mutations of the project or shared environments |
@@ -83,6 +84,19 @@ User-utterance interpretation rule:
83
84
 
84
85
  A single okstra run frequently spans 30–120 minutes with multi-minute silent windows while workers run; without progress signals the user cannot distinguish "still working" from "hung". Lead MUST emit a single short progress line at each checkpoint below — plain user-facing text in a separate brief message (not buried inside a tool call), one line per checkpoint, format: `PROGRESS: <phase-id> <verb-phrase>`. Emit the line raw — the literal `PROGRESS:` token must begin the line. Do NOT wrap it in inline-code backticks (`` `PROGRESS: ...` ``) or a ```` ``` ```` code fence; markdown wrapping is what the post-hoc conformance validator scrapes around, and raw emit keeps the signal unambiguous.
85
86
 
87
+ For an `implementation-planning` run whose run manifest declares `activityContractVersion: 1`, record every required activity boundary with `okstra agent-activity append` against the manifest-provided `leadEventsPath`. The ordering is fixed: the structured append succeeds first, the matching `PROGRESS:` line is emitted second, and the immediately following `ACTIVITY:` line projects the same structured fields into the conversation language. Do not reconstruct structured activity from conversation text. If the append fails, do not present that activity boundary as completed.
88
+
89
+ The live projection follows this shape:
90
+
91
+ ```text
92
+ PROGRESS: phase-4-dispatch worker=codex-worker model=gpt-5.6-sol
93
+ ACTIVITY: id=A-001 agent=codex-worker summary="Verify Stage Map paths and commands" items=P-Step-001,P-Step-002 result=runs/.../codex-worker-....md outcome=pending
94
+ ```
95
+
96
+ Use the exact CLI projection: `id`, `agent`, quoted `summary`, comma-joined `items` (`<none>` when empty), `result` (`<none>` when empty), and `outcome`. Only prose inside `summary` is localized to the conversation language. Required kinds are `worker-dispatched`, `worker-completed`, `verification-round-completed`, `self-fix-applied`, `user-decision-required`, and `user-decision-evaluated`.
97
+
98
+ **Enforcement:** `tests/contract/test_host_orchestration_rules.py` keeps this instruction on every lead path. `validators/validate_session_conformance.py` `_check_activity_contract` checks the structured event log and does not treat an `ACTIVITY:` conversation line as evidence.
99
+
86
100
  Required checkpoints:
87
101
 
88
102
  - `PROGRESS: phase-1-intake reading task bundle` — at the start of Phase 1, before issuing parallel Read calls.
@@ -106,7 +120,7 @@ Do NOT replace them with prose ("Now I'm starting Phase 2..."), do NOT skip a ch
106
120
 
107
121
  `okstra-run` surfaces these lines to the user directly; other launch paths persist them in the selected adapter's declared conformance evidence/event source for post-hoc retrieval.
108
122
 
109
- **Enforcement:** the Phase 7 validator (`validators/validate-run.py` → `validate_session_conformance.py`) reads the selected adapter's declared conformance evidence/event source within the run window and fails the run as `contract-violated` when a required checkpoint is missing — including the per-worker `phase-4-dispatch` / `phase-5-collect` lines (which must name each worker's role) and the `phase-batch-cleanup` lines that MUST precede the first `phase-5.5-convergence` round and the `phase-6-synthesis` report-writer dispatch. When the plan-body state file records two or more rounds, `_check_plan_verify_cleanup_checkpoints` additionally requires a `phase-5.5.9-plan-verify` line per round and a `phase-batch-cleanup` between consecutive rounds — that boundary sat outside both older checks, so a five-round self-fix loop left every round's verifiers holding their panes. `phase-7-teardown` and `complete` fire after validation and are not checked.
123
+ **Enforcement:** the Phase 7 validator (`validators/validate-run.py` → `validate_session_conformance.py`) reads the selected adapter's declared conformance evidence/event source within the run window and fails the run as `contract-violated` when a required checkpoint is missing — including the per-worker `phase-4-dispatch` / `phase-5-collect` lines (which must name each worker's role) and the `phase-batch-cleanup` lines that MUST precede the first `phase-5.5-convergence` round and the `phase-6-synthesis` report-writer dispatch. When the plan-body state file records two or more rounds, `_check_plan_verify_cleanup_checkpoints` additionally requires a `phase-5.5.9-plan-verify` line per round and a `phase-batch-cleanup` between consecutive rounds. For activity-contract-v1 planning, `_check_activity_contract` validates the structured worker pairs, verification and self-fix counts, user-decision references, and `A-NNN` ordering. `phase-7-teardown` and `complete` fire after validation and are not checked.
110
124
 
111
125
  ## User confirmation before an approval blocker (BLOCKING)
112
126
 
@@ -118,9 +132,22 @@ The sequence is fixed:
118
132
 
119
133
  1. Emit `PROGRESS: user-confirm <C-NNN> <the question, one line>` with the id the row would carry.
120
134
  2. Ask in plain user-facing text: what is undecided, the options with their consequences, and which one you recommend. One question at a time.
121
- 3. On an answer — record it in the row's `userInput`, set `status: answered` and `userConfirmation: asked-and-answered`, apply it, and **keep going in this run**. An answered question is not a blocker, and a run that stops anyway wastes the answer it just received.
135
+ 3. On an answer — record the raw text in the row's `userInput`, set `status: answered` and `userConfirmation: asked-and-answered`, and apply the selected disposition in this run.
122
136
  4. Only when asking fails does the row stay open: `asked-awaiting` when the user has not answered, `deferred-no-interactive-session` when this run has no user to ask.
123
137
 
138
+ For activity-contract-v1 `implementation-planning`, every approval row carries `approvalContext`. Classify a user-owned selection as `user-decision`, a surviving non-correctness majority disagreement as `noncritical-dissent`, and a cited path/symbol mismatch, `P-Req-*` coverage mismatch, or independent Requirement Coverage blocker as `correctness-critical`. `select` is limited to `user-decision`, `accept-risk` is limited to `noncritical-dissent`, and `request-revision` / `reject` are available to all three classifications. `correctness-critical` never offers or records `accept-risk`. **Enforced:** `validators/validate-run.py` `_validate_approval_context` recomputes the classification, disposition allowlist, activity references, and resolved-state requirements.
139
+
140
+ The approval state transitions are fixed:
141
+
142
+ - `open → answered` when the raw user response is recorded
143
+ - `answered → resolved` only after the selected disposition is applied and its checks pass
144
+ - `answered → open` when application or checking fails
145
+ - `open → obsolete` only when a plan change removes the question
146
+
147
+ `open` and `answered` continue to block approval; only `resolved` and `obsolete` are non-blocking. `user-decision` resolves after the choice is applied and structure / extraction / Requirement Coverage checks pass. `noncritical-dissent` resolves only after an explicit `accept-risk` with non-empty user text and activity-backed checks. `correctness-critical` resolves only after the correction's targeted re-verification records `AGREE` or an acceptable `SUPPLEMENT` for every linked item and no independent coverage blocker remains. A user-directed correction does not consume the automatic self-fix limit, and a verification failure after that correction does not restart the automatic loop.
148
+
149
+ When a terminal row preserves a pre-correction dissent classification, keep the superseded votes in `state/plan-body-verification-implementation-planning-<seq>.json`; the validator recomputes the historical class from those votes and never trusts `approvalContext.classification` alone. Each cited `user-decision-required` and `user-decision-evaluated` activity records the row's exact `C-NNN` in `evidenceRefs` and covers every `approvalContext.planItemIds` value; an evaluated activity also records check evidence beyond the `C-NNN` itself. A corrected coverage-only blocker keeps its `C-NNN` in the non-blocking Requirement Coverage row's `decisionRefs`, and the state-sidecar plan item without a historical blocking dissent that participated in the `coverage-gap` round keeps the same `C-NNN` in `clarificationId`; a run-wide `coverage-gap` without that item-level link is not evidence for the row. An `obsolete` row is invalid while its disagreement or coverage blocker remains active in the current plan. **Enforced:** `validators/validate-run.py` `_read_approval_history`, `_activity_matches_approval_context`, `_historical_coverage_clarification_ids`, and `_validate_approval_context`.
150
+
124
151
  **Predicting the blocker is not the same as raising it.** A lead that says "this will likely become an approval blocker; I will ask at that point" has already reached the moment — ask then, in that message. One run announced exactly that, never asked, wrote the row anyway, and then spent its entire self-fix budget on a gate no round could clear, because the user had already answered the question before the run started.
125
152
 
126
153
  **`lead-directed` blockers cannot be deferred.** When the item is the lead's own judgment rather than a worker's finding, and this run has nobody to ask, the row is not the outlet — record a Working Assumption in `## 5. Missing Information and Risks` naming the assumption the plan proceeds under, exactly as a surviving planner-fixable item does, and let the plan proceed. Blocking a plan on the lead's own judgment in a run where that judgment cannot be put to the user only moves the work to a re-run.
@@ -162,7 +189,7 @@ Executor is chosen at run-prep time via `--executor <claude|codex|antigravity>`
162
189
 
163
190
  `okstra-ctl` provisions dedicated `git worktree`s at run-prep time. Lead, the Executor, and every verifier MUST treat the provisioned worktree as the canonical working directory regardless of task-type.
164
191
 
165
- - **Task-key worktree (non-`implementation` phases):** `requirements-discovery`, `error-analysis`, and `implementation-planning` share one worktree per task-key so phase N inherits the working-tree state phase N-1 left behind. Location: `~/.okstra/worktrees/<project-id>/<task-group-segment>/<task-id-segment>/` (override `OKSTRA_HOME` only for tests). All segments are sanitised — `/`, `:`, and other special chars collapse to `-`.
192
+ - **Task-key worktree (non-`implementation` phases):** `requirements-discovery`, `error-analysis`, `implementation-option-selection`, and `implementation-planning` share one worktree per task-key so phase N inherits the working-tree state phase N-1 left behind. Location: `~/.okstra/worktrees/<project-id>/<task-group-segment>/<task-id-segment>/` (override `OKSTRA_HOME` only for tests). All segments are sanitised — `/`, `:`, and other special chars collapse to `-`.
166
193
  - **Stage worktree (`implementation`):** stage-isolated — one run = one stage, each in its own worktree at `.../<task-id-segment>/stage-<N>/` on its own branch. Single-stage `final-verification` (`--stage <N>`) reuses that stage worktree read-only; whole-task `final-verification` operates on the task-key worktree.
167
194
  - Branch: `<work-category-namespace>/<task-id-segment>` (e.g. `feature/dev-9436`, `fix/dev-7311`); a stage worktree appends `-s<N>` (e.g. `feature/dev-9436-s2`). The task-key worktree is branched from the user-chosen `--base-ref` (default: `HEAD` of the repo's **main** worktree) at the first phase's prep time; a stage worktree's base is resolved from its `depends-on` anchors at prep time. The resolved base SHA is recorded in `EXECUTOR_WORKTREE_BASE_REF`.
168
195
  - A global registry at `~/.okstra/worktrees/registry.json` (flock-guarded) reserves both task-keys and stage-keys (`<task-key>#stage-<N>`), mapping each to its path + branch, and prevents concurrent runs from colliding. Branch names are globally unique on this machine.
@@ -206,7 +233,7 @@ The `implementation` profile's thin core (`prompts/profiles/implementation.md`)
206
233
 
207
234
  The guard is not satisfied by memory from a prior run — each implementation run re-reads the sidecar fresh, since `okstra install` may have updated it between runs.
208
235
 
209
- This pattern is implementation-only. Other profiles (`requirements-discovery`, `error-analysis`, `implementation-planning`, `final-verification`, `release-handoff`) load their whole profile body at Phase 1 as before — they are short enough not to benefit from a split.
236
+ This pattern is implementation-only. Other profiles (`requirements-discovery`, `error-analysis`, `implementation-option-selection`, `implementation-planning`, `final-verification`, `release-handoff`) load their whole profile body at Phase 1 as before — they are short enough not to benefit from a split.
210
237
 
211
238
  Extract from the compact intake files: task key, task type, work category, workflow lifecycle snapshot, selected worker roster, assigned models, worker result paths, worker prompt history paths, current run prompt directory, final report path, final status path, validator path, resume helper path, config-file references, deployment-manifest references, and their expected values or invariants.
212
239
 
@@ -304,7 +331,7 @@ Convergence is enabled by default. Configure via task-manifest.json:
304
331
  - `convergence.enabled`: true/false (default: true)
305
332
  - `convergence.maxRounds`: 1–3 — **phase-aware default**: `1` for `requirements-discovery`, `2` for all other task types
306
333
  - `convergence.verificationMode`: `"lightweight"` | `"full-reanalysis"` (default: `"lightweight"`; the adversarial phases below force `"full-reanalysis"`)
307
- - `convergence.adversarial`: true/false — **phase-aware default**: `true` for `requirements-discovery` / `error-analysis` / `implementation-planning` / `project-analysis` / `feature-analysis` / `change-impact-analysis`, `false` otherwise. When `true`, Phase 5.5 runs in adversarial mode (verifiers refute findings; burden of proof on the claim). See [convergence](./convergence.md) "Adversarial Verification Mode".
334
+ - `convergence.adversarial`: true/false — **phase-aware default**: `true` for `requirements-discovery` / `error-analysis` / `implementation-option-selection` / `implementation-planning` / `project-analysis` / `feature-analysis` / `change-impact-analysis`, `false` otherwise. When `true`, Phase 5.5 runs in adversarial mode (verifiers refute findings; burden of proof on the claim). See [convergence](./convergence.md) "Adversarial Verification Mode".
308
335
 
309
336
  When `task-manifest.json` does not set `convergence.maxRounds`, lead MUST resolve the effective value via the phase-aware default above before entering Phase 5.5 and put it in the grouped input at `config.effectiveMaxRounds`.
310
337
 
@@ -351,7 +378,7 @@ After the Report writer worker draft is reviewed (or after the lead-authored fal
351
378
 
352
379
  This is a Phase 6 sub-step — it does NOT introduce a new top-level lifecycle phase; the lead operating-phase model (Phase 1 Intake → Phase 7 Persist, labels in the "Quick Reference" table above as the single source of truth) is preserved. The round's outcome is read from the final report's `### 5.5.9 Plan Body Verification` section and `implementationPlanning.planBodyVerification` in its data.json — it is not a separate lifecycle phase identifier.
353
380
 
354
- **REQUIRED RESOURCE:** Read [plan-body-verification](./plan-body-verification.md) for the round protocol, plan-item ID scheme (`P-Opt-*` / `P-Step-*` / `P-Dep-*` / `P-Val-*` / `P-Rb-*` / `P-Req-*` / `P-Prep-*`), verdict semantics (`AGREE` / `DISAGREE(a-f)` / `SUPPLEMENT`), classification rules, gate-result resolution, and the state-file schema at `runs/<task-type>/state/plan-body-verification.json`.
381
+ **REQUIRED RESOURCE:** Read [plan-body-verification](./plan-body-verification.md) for the round protocol, plan-item ID scheme (`P-Dir-1` for selected-direction; `P-Opt-*` for legacy candidate comparison; then `P-Step-*` / `P-Dep-*` / `P-Val-*` / `P-Rb-*` / `P-Req-*` / `P-Prep-*`), verdict semantics (`AGREE` / `DISAGREE(a-f)` / `SUPPLEMENT`), classification rules, gate-result resolution, and the state-file schema at `runs/<task-type>/state/plan-body-verification.json`. For `P-Dir-1`, compare `directionRealization` with `selectedDirectionRef` and its snapshot: verify the core mechanism, architecture boundaries, planning invariants, and any hidden direction change.
355
382
 
356
383
  Distinct from Phase 5.5 finding convergence:
357
384
 
@@ -361,6 +388,8 @@ Distinct from Phase 5.5 finding convergence:
361
388
 
362
389
  Lead's responsibilities in this sub-step (in order):
363
390
 
391
+ For a new `implementation-planning` run, the fixed order is initial verification → one planner self-fix → targeted re-verification → user gate. The initial verification is round 1 and the targeted re-verification is round 2. A second automatic self-fix is a contract violation.
392
+
364
393
  1. Build the queue with `okstra plan-items extract --data <data.json> --output <state>/plan-items-....json`, place the persisted `items[]` verbatim in every verifier prompt, then run `okstra plan-items validate --data <data.json> --items <state>/plan-items-....json`. The lead MUST NOT summarise, select, omit, reorder, or renumber the queue. Each prompt uses the compact `subject` plus the lossless `payload`, and asks every item:
365
394
 
366
395
  ```text
@@ -371,7 +400,7 @@ Lead's responsibilities in this sub-step (in order):
371
400
  An `AGREE` response records the considered counterexample and exclusion reason in its note; unverified external material is `verification-error`, not `DISAGREE`.
372
401
  2. Dispatch a single plan-body reverify round to every analyser worker in the roster (`claude`, `codex`, and `antigravity` when opted in). `Report writer worker` is NOT a participant in this round.
373
402
  3. Aggregate verdicts and resolve the gate result to one of `passed` / `passed-with-dissent` / `blocked-by-disagreement` / `aborted-non-result`.
374
- 4. Write `runs/<task-type>/state/plan-body-verification.json` (schema in the plan-body-verification contract), appending one `roundHistory[]` entry per round — including every self-fix re-verification round, since data.json keeps only the final verdicts.
403
+ 4. Write `runs/<task-type>/state/plan-body-verification.json` (schema in the plan-body-verification contract), appending round 1 and, if the one automatic rewrite ran, round 2 to `roundHistory[]`; data.json keeps only the final verdicts.
375
404
  5. Populate `implementationPlanning.planBodyVerification` in data.json with round count, gate result, per-item verdicts, and dissent log. The AI handoff task-deliverable block carries this structure without a second prose rendering.
376
405
  6. For every `majority-disagree` plan item, append one `clarificationItems[]` row with `blocks=approval` and the 1:1 ID match in the verdict classification (`majority-disagree → C-<N>`). Do not create a parallel open-questions structure.
377
406
  7. Publish the YAML frontmatter `approved:` field as `false`. There is no in-body `- [ ] Approved` marker line — approval lives only in the frontmatter (see [plan-body-verification](./plan-body-verification.md) §"Round protocol" step 9). The user may flip it to `true` only when the gate is `passed` or `passed-with-dissent`. **Enforced:** `validators/validate-run.py` `validate_phase_boundary` fails a report shipping `approved: true` under `blocked-by-disagreement` / `aborted-non-result`, and run-prep (`scripts/okstra_ctl/run.py` `_validate_approved_plan`) fail-closes the same case. Manually flipping a blocked gate to passing is a contract violation.
@@ -384,14 +413,10 @@ The detailed persistence checklist and the BLOCKING token-usage collector invoca
384
413
 
385
414
  Order of operations:
386
415
 
387
- 0. **Translation (conditional).** Read `meta.reportLanguage` from the final-report data.json. When it is `en`, skip this step entirely — there is nothing to translate and no worker to pay for. When it is anything else, dispatch `Translator worker` with `**Report Language:**` set to that value and `**Result Path:**` set to `runs/<task-type>/reports/final-report-<task-type>-<seq>.i18n.<lang>.json`. The worker builds its own work list with `okstra report-translate extract` and gates itself with `okstra report-translate check`; Lead verifies the sidecar exists before continuing. The data.json stays English — the sidecar is presentation, overlaid by `render-views` in step 2, and a missing one degrades to an English HTML rather than failing the run.
388
- 1. Run the token-usage collector with `--substitute-data`. The final-report data.json MUST already exist; the collector populates its token / cost cells and re-renders the markdown sibling.
389
- 2. Verify both the final-report data.json and rendered markdown exist at the expected paths. When `Report writer worker` is in the roster, that worker authored the data.json + invoked the renderer in Phase 6 — Lead's job here is to **verify** schema validation passes and structure matches the template output.
390
- 3. Update team-state artifact (preserve usage fields written by the script).
391
- 4. Update run manifest.
392
- 5. Update `task-manifest.json`, including lifecycle fields (work category, phase states, next recommended phase, approval markers, safe-resume checkpoint).
393
- 6. Update `task-index.md`.
394
- 7. Write final status file if expected.
416
+ 1. Run `okstra agent-activity project --project-root <root> --run-manifest <path> --data <data.json>`. This deterministically projects the canonical lead-events activity rows before any prose inspection.
417
+ 2. Run `okstra report-translate check-source <data.json>` even when `meta.reportLanguage` is `en`.
418
+ 3. When `meta.reportLanguage` is not `en`, dispatch the translator worker. The worker builds its work list with `okstra report-translate extract`, writes `final-report-<task-type>-<seq>.i18n.<lang>.json`, and gates it with `okstra report-translate check`.
419
+ 4. Run `okstra report-finalize ...`. This command owns token substitution, view rendering, follow-up persistence, and validation in their contractual order.
395
420
 
396
421
  Keep the assigned worker prompt history paths stable in `team-state`, `run-manifest`, and `task-manifest`. Do not rewrite prompt artifacts to `/tmp` or omit prompt metadata for attempted workers.
397
422
 
@@ -443,4 +468,4 @@ After persistence, reply briefly in the resolved Report Language with: completio
443
468
  | Waiting silently after `dispatch_worker` returns without a completed worker artifact | A dispatch acknowledgement is not completion — call `await_workers` and enforce the selected adapter's liveness policy |
444
469
  | Re-sending a finding absent from the persisted round plan | Dispatch exactly the engine-returned `findingIds`; see [convergence](./convergence.md) "Re-verification Dispatch" |
445
470
  | Aggregating a `timeout`/`error` reverify dispatch as `DISAGREE` | Put the terminal outcome in round results; `apply-round` records `verification-error`. See [convergence](./convergence.md) "Worker failure handling in reverify" |
446
- | Skipping `--substitute-data` in the Phase 7 collector run | Always pass the flag — see [report-writer](./report-writer.md) "Phase 7 token-usage collector" |
471
+ | Bypassing `report-finalize` and running its Phase 7 steps manually | Run `okstra report-finalize ...`; it owns token substitution and the remaining persistence order. |
@@ -44,7 +44,7 @@ Plan-body verification is configured under `convergence.planBodyVerification` in
44
44
  |---------|---------|-------------|
45
45
  | `enabled` | `true` | If `false`, the round is skipped and the approval gate is not blocked by this round (legacy behaviour). |
46
46
  | `maxRounds` | `1` | Upper bound. Plan-body verification is consistency / completeness checking, not fact checking — additional rounds rarely help. Range 1–3. |
47
- | `selfFixMaxRounds` | `3` | Upper bound on the report-writer self-fix loop (§"Round protocol" step 7). Range 1–5. The loop also stops early on no-progress, so this is a ceiling, not a target. |
47
+ | `selfFixMaxRounds` | `1` | One report-writer rewrite at most. The initial verification is round 1; targeted re-verification is round 2 after that rewrite. |
48
48
  | `gating` | `true` | If `true` (default), `majority-disagree` blocks approval. If `false`, the round is advisory-only and never blocks approval. |
49
49
 
50
50
  Default values are emitted into the manifest by `scripts/okstra_ctl/render.py` (`_build_convergence_block`). The ctx knob `OKSTRA_PLAN_VERIFICATION=false` flips `planBodyVerification.enabled` to false.
@@ -67,11 +67,26 @@ as its heading, but it MUST include the lossless `payload` for the item's eviden
67
67
  judgement. The final `validate` command confirms that the persisted queue exactly matches
68
68
  the current draft before verdict aggregation.
69
69
 
70
- The deterministic extractor assigns the following prefixes:
70
+ The deterministic extractor assigns one contract-specific direction prefix, followed by the shared execution prefixes.
71
+
72
+ ### Legacy candidate-comparison branch
73
+
74
+ | ID | Source | Payload |
75
+ |---|---|---|
76
+ | `P-Opt-<N>` | `4.5.1 Option Candidates` | one Option (its File Structure list + interfaces + blast radius); verify its trade-off claims and consistency with the recommended option |
77
+
78
+ ### Selected-direction branch
79
+
80
+ | ID | Source | Payload |
81
+ |---|---|---|
82
+ | `P-Dir-1` | `implementationPlanning.directionRealization` | exactly one selected-direction realization; compare it with `selectedDirectionRef` and the byte-verified snapshot |
83
+
84
+ `P-Dir-1` verifies the core mechanism, architecture boundaries, planning invariants, and any hidden direction change. An AGREE verdict means `directionRealization` preserves those properties from the snapshot named by `selectedDirectionRef`; it does not re-score candidates or recommend another direction. A required direction change is a `direction-invalidated` result, not a planner rewrite.
85
+
86
+ ### Shared execution items
71
87
 
72
88
  | Prefix | Source sub-section | One row per |
73
89
  |--------|--------------------|-------------|
74
- | `P-Opt-<N>` | `4.5.1 Option Candidates` | one Option (its File Structure list + interfaces + blast radius) |
75
90
  | `P-Step-<N>` | `4.5.4 Stepwise Execution Order` | one step (path + command + success signal) |
76
91
  | `P-Dep-<N>` | `4.5.5 Dependency / Migration Risk` | one dependency row |
77
92
  | `P-Val-<N>` | `4.5.6 Validation Checklist` | one checklist item |
@@ -80,7 +95,7 @@ The deterministic extractor assigns the following prefixes:
80
95
  | `P-Prep-S<stage>-<kind>` | Stage `designSurfaceCoverage` + `5.5.10 Implementation Design Preparation` | exactly one detector-produced `(stage, kind)` |
81
96
  | `P-Var-<N>` | `5.5.11 Variation-Point Analysis` | one variation point (its `behavior` + `extractionDecision`), or a lone `P-Var-0` when the plan declares no variation point |
82
97
 
83
- `4.5.2 Trade-off Matrix` and `4.5.3 Recommended Option` are NOT extracted as standalone plan items — the trade-off matrix is evaluated implicitly through each option's `P-Opt-*` verification, and the recommended option is one of those `P-Opt-*` rows.
98
+ For legacy candidate-comparison plans, `4.5.2 Trade-off Matrix` and `4.5.3 Recommended Option` are NOT extracted as standalone plan items — the trade-off matrix is evaluated implicitly through each option's `P-Opt-*` verification, and the recommended option is one of those `P-Opt-*` rows. Selected-direction plans contain neither section and use only `P-Dir-1` for direction preservation.
84
99
 
85
100
  Each plan item inherits the `[TICKETID: ...]` tag of its source section (per the standard ticket-tagging contract).
86
101
 
@@ -123,6 +138,8 @@ DISAGREE on a `P-Var-*` item means one of:
123
138
 
124
139
  The hexagonal rule that an extracted point must declare `interfaceKind: "port"` is already machine-checked by `validators/validate-run.py` `_validate_variation_point_analysis` (it fires only for a project whose `architecture.style` is `hexagonal`). Do not re-run that mechanical check as a verdict; spend the judgement on placement and semantics instead — a point extracted as a port whose domain rule leaked into the adapter passes the validator and is still wrong.
125
140
 
141
+ `P-Dir-1` carries the same YAGNI judgement as the legacy option item, but its comparison source is the selected-direction snapshot rather than a trade-off matrix. A new abstraction, configuration knob, widened interface, file, or stage with no original-requirement link is a hidden direction change and receives `DISAGREE(e)` on `P-Dir-1`.
142
+
126
143
  `P-Opt-<N>` carries the **YAGNI judgement** and is majority-gated for the same reason as `P-Var-*`: whether an abstraction serves the stated requirement or only a forecast is a judgement about the design, not a contradiction between two spelled-out references. Raise it as `DISAGREE(e)` — an option that carries an abstraction, parameter, or configuration knob no Requirement Coverage row demands contradicts the trade-off matrix that scored it, because the complexity the matrix priced is not the complexity the option actually buys. DISAGREE on a `P-Opt-*` item under this rule means one of:
127
144
 
128
145
  - **an abstraction nobody asked for** — a helper module, strategy / factory, indirection layer, or interface whose only justification in the plan is a caller no requirement names. A second implementation already on the table is `P-Var-*` territory and is the opposite defect: do not raise both on the same behavior;
@@ -237,35 +254,44 @@ round before any host or provider process starts.
237
254
 
238
255
  **How the corrective round is recorded.** The first prompt was dispatched, so it is immutable — `--replace-undispatched` refuses it, correctly. Materialize the correction under a NEW `--invocation-id` and a new prompt path. Before linking its result, retire the first attempt's link: `okstra agent-prompt reject-result --run-manifest <path> --dispatch-id <first dispatch id> --superseded-by <corrective dispatch id> --reason "<what was wrong with the returned result>"`. Without that step the corrective `link-result` fails with `agent result is already linked to another dispatch`, which is how a worker that ran for twenty minutes and wrote a good result ends up unrecordable. Nothing is deleted: the rejected link stays in `agentResultLinks` carrying `supersededBy` and `rejectionReason`, so the ledger shows both attempts and why the second exists.
239
256
 
240
- Then lead writes `runs/<task-type>/state/plan-body-verification-<task-type>-<seq>.json` (schema below), **appending this round** — one new `roundHistory[]` entry plus this round's votes on each verified item's `planItems[].rounds[]`. The file accumulates across rounds; it is never truncated to the latest one. Lead then populates `### 5.5.9 Plan Body Verification` in the final report's data.json (`implementationPlanning.planBodyVerification`, schema `schemas/final-report-v1.0.schema.json`; template at `templates/reports/final-report.template.md`). The §5.5.9 body is **grouped by plan item**: `planItems[]`, each carrying its `id`, its plain-language `subject` (rendered as the item heading), an optional `sourceSection`, an optional `clarificationId` (the `C-<N>` this item blocks on when `majority-disagree`), and a `verdicts[]` list (`worker / verdict / breakageKind / note`) — one verdict row per worker under that item. The renderer prints three fixed legends (gate values, verdict tokens, breakage kinds a–f) so the reader can decode every cell without opening this spec. The older flat `#### Verdict details` table (`Plan item / Worker / …`, one row per plan-item × worker pair) is superseded by the grouped layout — it hid *what* each vote was about behind a bare `P-*` ID; the subject heading is the fix. The validator's `Plan Body Verification` + `Gate result:` substring checks still gate this section.
241
- 7. **Self-fix loop (up to `selfFixMaxRounds`, targeting planner-fixable defects).** After aggregation, while at least one `majority-disagree` item has a majority of its `DISAGREE` verdicts at `fixability == planner-fixable`, lead runs self-fix rounds **before** promoting anything to the user:
257
+ Then lead writes `runs/<task-type>/state/plan-body-verification-<task-type>-<seq>.json` (schema below), **appending this round** — one new `roundHistory[]` entry plus this round's votes on each verified item's `planItems[].rounds[]`. The file accumulates across rounds; it is never truncated to the latest one. After `okstra plan-verify` exits 0, lead sets that new round's `completedAt` to the current ISO 8601 UTC time exactly once; a prior round's `completedAt` is immutable. Lead then populates `### 5.5.9 Plan Body Verification` in the final report's data.json (`implementationPlanning.planBodyVerification`, schema `schemas/final-report-v1.0.schema.json`; template at `templates/reports/final-report.template.md`). The §5.5.9 body is **grouped by plan item**: `planItems[]`, each carrying its `id`, its plain-language `subject` (rendered as the item heading), an optional `sourceSection`, an optional `clarificationId` (the `C-<N>` this item blocks on when `majority-disagree`), and a `verdicts[]` list (`worker / verdict / breakageKind / note`) — one verdict row per worker under that item. The renderer prints three fixed legends (gate values, verdict tokens, breakage kinds a–f) so the reader can decode every cell without opening this spec. The older flat `#### Verdict details` table (`Plan item / Worker / …`, one row per plan-item × worker pair) is superseded by the grouped layout — it hid *what* each vote was about behind a bare `P-*` ID; the subject heading is the fix. The validator's `Plan Body Verification` + `Gate result:` substring checks still gate this section.
258
+ 7. **Self-fix loop (one rewrite, targeting planner-fixable defects).** After round 1, lead may run one report-writer rewrite when at least one `majority-disagree` item has a majority of its `DISAGREE` verdicts at `fixability == planner-fixable`. The targeted re-verification after that rewrite is round 2. After round 2, stop automatic self-fix regardless of outcome. Classify every remaining item as `user-decision`, `noncritical-dissent`, or `correctness-critical`. A second automatic self-fix is a contract violation. The fixed order is initial verification → one planner self-fix → targeted re-verification → user gate.
242
259
  - **Group the targets by cause before instructing (BLOCKING).** Blocked items are usually several derivatives of one defect — one constant declared twice, one responsibility given two owners — and the coverage rows that cite them fail as a consequence, not independently. Lead MUST partition this round's targets into cause groups and instruct each group as **"remove this cause"**, naming the derivatives it accounts for. **Handing report-writer a bare item list is forbidden**: patched one at a time, each correction leaves the sibling sections still asserting the old value, so the next round re-finds the same family and the budget drains without converging. Record the partition in `planBodyVerification.selfFixGroups[]` (`round`, `causeSummary`, `itemIds`). One group per item is a legitimate outcome only when the items genuinely share no cause — recorded that way, it is a visible diagnosis rather than a skipped one. **Enforced:** `validators/validate-run.py` `_validate_self_fix_grouping` requires the partition, ties `selfFixRoundsApplied` to the highest recorded round, and fails any corrected item that belongs to no group.
243
260
  - lead instructs report-writer to rewrite the items in each cause group (NOT a full draft regeneration; procedure in [report-writer](./report-writer.md) §"Self-fix rewrite").
244
261
  - missing or weak `P-Prep-*` contracts are repaired by adding kind-specific inline detail or an AI-prepared PREP item with a concrete proposal. Facts that require user or external authority remain `blocked` and keep their request material; never invent those facts during self-fix.
245
262
  - **Drop plan items whose element the round deleted.** A self-fix rewrite may remove a plan element (a validation check, a rollback row). `P-*` ids are positional, so a deletion shifts every later row and silently re-points surviving verdicts at their neighbours — and a verdict recorded against a removed element keeps blocking a gate while being unfindable in the plan, so reading the plan never reveals the cause. After each round, re-extract plan items with `okstra plan-items extract` and re-verify any item whose `subject` no longer matches; never carry the old vote forward across a shift. **Enforced:** `validators/validate-run.py` `_validate_verdicts_match_current_subjects` (re-pointing) and `_validate_plan_item_extraction_completeness` (dangling ids).
246
263
  - **Classify each cause group before instructing it (BLOCKING).** A group is either an *authoring* defect — the plan says something wrong, incomplete, or self-contradictory, which self-fix owns — or a *citation* defect, where the plan points at an analysis artifact incorrectly. Only the first is self-fix work. For the second the finding already exists and already went through convergence, so the fix is to re-cite the converged artifact; instructing report-writer to re-derive the fact means the author reads the source material and produces a **finding that never went through convergence**, which the plan then carries as if it had. That is the role boundary the lead contract draws ("keep analysis, execution, verification, and report authoring responsibilities distinct; return defects to the role that owns them"), and report-writer is authoring-only by its own contract. `P-Req-*` items with breakage kind `f` are where this goes wrong most often: the question is usually whether a coverage row points correctly at something already measured, not whether the measurement is right. State the classification in the group's instruction so the author knows which of the two it is being asked to do.
247
- - **A verdict older than the last self-fix is not a verdict (BLOCKING).** Rounds interleave with rewrites — round 1, self-fix 1, round 2, self-fix 2 — so a verdict cast in round R judged the text as it stood after self-fix R-1. Once self-fix R runs, that judgement is about a plan that no longer exists. `--round <N>` on `apply-verdicts` stamps each row, and `validators/validate-run.py` `_validate_verdict_rounds_outlive_self_fix` fails any non-carried item whose verdict round is at or before `selfFixRoundsApplied`. This is why "adjacent items the rewrite touched" is not sufficient on its own: adjacency is judged from `subject` changes, and the observed failure was items whose own subject never moved while the stage they point at was rewritten under them. On one run the gate read `passed-with-dissent` with zero blockers and a single re-run flipped 3 of 27 items to `majority-disagree`, all correctness-critical. Before declaring the gate, every item still holding a pre-self-fix verdict MUST be re-verified in a round after the last rewrite.
264
+ - **A verdict older than the last self-fix is not a verdict (BLOCKING).** A verdict cast in round 1 judged the text before the only automatic rewrite. Once that rewrite runs, the judgement is about a plan that no longer exists. `--round <N>` on `apply-verdicts` stamps each row, and `validators/validate-run.py` `_validate_verdict_rounds_outlive_self_fix` fails any non-carried item whose verdict round is at or before `selfFixRoundsApplied`. Before declaring the gate, every item still holding a pre-self-fix verdict MUST be re-verified in round 2.
248
265
  - lead re-runs plan-body verification (focused on the corrected items + adjacent items the rewrite touched, plus any `needs-reverify` items whose peer failed to vote last round). After re-verification, overwrite `planItems[].verdicts` with the new verdicts. **The round's verdicts MUST be transcribed into `planBodyVerification.planItems[].verdicts` in the final report's data.json before the gate is declared** — the gate is re-derived from that table, so declaring a gate over an empty one leaves it unauditable. **Enforced:** `_validate_round_recorded_verdicts`. Transcribe with `okstra plan-items collect-verdicts --result <worker>=<path> … --items <plan-items.json> --output <verdicts.json>` then `okstra plan-items apply-verdicts --data <data.json> --verdicts <verdicts.json> --round <N>`, never with a per-round script: the CLI reads the response shape this section fixes and **fails** on an assigned item the worker left unanswered, on a verdict for an item outside the queue, and on a `DISAGREE` with no breakage kind. A hand-written regex reports none of those — it drops them, and the round is then scored on a table that silently does not match the queue.
249
266
  - for an item whose `majority-disagree` was resolved by self-fix, record `self-fixed in round <N>: <what was fixed>` in `planItems[].selfFixNote`. A resolved item does not create a clarification.
250
267
  - **Each round is a worker batch.** Before dispatching round N ≥ 2, reclaim the previous round's completed verifiers exactly as at any other batch boundary ([okstra-lead-contract](./okstra-lead-contract.md) "Run-scoped worker-resource lifecycle") and emit `PROGRESS: phase-batch-cleanup panes=<n>`, then announce the round with `PROGRESS: phase-5.5.9-plan-verify round=<N> items=<count>`. Saying a round will "reuse" the previous verifiers and then dispatching under fresh names leaves every prior round holding its panes — five rounds of that is what exhausts the pane budget and blocks the next dispatch. **Enforced:** `validators/validate_session_conformance.py` `_check_plan_verify_cleanup_checkpoints` requires both lines once the state file records two or more rounds.
251
- - **Round completion.** A round is complete only after the renderer has run on the corrected data.json, lead has appended the round to the state file per step 6, lead has reconciled instructed groups against applied corrections — every `itemIds` entry either carries a `selfFixNote` or is still recorded as broken — and **`okstra plan-verify --report <report>` exits 0** (step 5). A round left with a non-zero exit carries its defect into the next round's inputs, which is how a mis-scored gate survives a whole self-fix budget. A round that was instructed but never rendered has not happened, and counting it inflates the budget that gates promotion. The state-file append is not optional bookkeeping: the next re-verification overwrites data.json's `planItems[].verdicts`, so a round that never reached `roundHistory[]` leaves no record anywhere of what it blocked on — which is the whole reason this file exists. **Enforced:** `validators/validate-run.py` `_validate_plan_body_state_rounds` requires one `roundHistory[]` entry per round `1..roundCount`, each carrying its own `gateResult` and cited by at least one item's `rounds[]`, and requires the file's `selfFixRoundsApplied` to match the report's.
252
- - **Loop termination.** Lead — not the report-writer worker — records the round count in `planBodyVerification.selfFixRoundsApplied` at each round's end, and why the loop stopped in `planBodyVerification.selfFixStopReason`. The count must equal the highest `round` in `selfFixGroups[]`, so it is derivable from recorded work rather than self-reported:
268
+ - **Round completion.** A round is complete only after the renderer has run on the corrected data.json, lead has appended the round to the state file per step 6, lead has reconciled instructed groups against applied corrections — every `itemIds` entry either carries a `selfFixNote` or is still recorded as broken — **`okstra plan-verify --report <report>` exits 0** (step 5), and lead has then set that round's immutable `completedAt`. A round left with a non-zero exit carries its defect into the next round's inputs, which is how a mis-scored gate survives a whole self-fix budget. A round that was instructed but never rendered has not happened, and counting it inflates the budget that gates promotion. The state-file append is not optional bookkeeping: the next re-verification overwrites data.json's `planItems[].verdicts`, so a round that never reached `roundHistory[]` leaves no record anywhere of what it blocked on — which is the whole reason this file exists. **Enforced:** `validators/validate-run.py` `_validate_plan_body_state_rounds` requires one `roundHistory[]` entry per round `1..roundCount`, each carrying its own `gateResult` and cited by at least one item's `rounds[]`, and requires the file's `selfFixRoundsApplied` to match the report's. For a resolved correctness-critical user response, `_validate_target_round_causality` additionally requires the cited round's `completedAt` to be after every linked canonical `user-decision-required` event and no later than every linked canonical `user-decision-evaluated` event.
269
+ - **Loop termination.** Lead — not the report-writer worker — records the round count in `planBodyVerification.selfFixRoundsApplied` at each round's end, and why the loop stopped in `planBodyVerification.selfFixStopReason`. The count must equal the highest `round` in `selfFixGroups[]`, so it is derivable from recorded work rather than self-reported. A user-directed correction does not consume the automatic self-fix limit, and a verification failure after that correction does not restart the automatic loop:
253
270
  - `all-resolved` — no planner-fixable `majority-disagree` item remains. Exit.
254
271
  - `no-progress` — the round resolved **zero** planner-fixable items relative to the previous round. Exit even with budget left: the same rewrite would repeat. Newly *introduced* defects count against progress, so a rewrite that trades one defect for another stops the loop rather than churning.
255
272
  - `max-rounds-reached` — `selfFixRoundsApplied == selfFixMaxRounds`. Exit.
256
- - `cause-group-recurrence` — this round's `itemIds` are a subset of what the *previous* round left unresolved (its targets minus the ones that earned a `selfFixNote` for that round). Exit even with budget left. `no-progress` counts resolutions, which a round can claim while its own correction plants the next round's defect; repeating the previous round's unresolved remainder is the earlier signal for the same dead end. Narrowing onto genuinely-open items is progress and does NOT fire this. **Enforced (advisory):** `validators/validate-run.py` `_detect_self_fix_recurrence` reports the shape as a warning — it does not fail the run, because the boundary between narrowing and re-digging is not settled.
273
+ - `cause-group-recurrence` — legacy read-only value for reports produced before activity contract v1. A new activity-contract-v1 run MUST NOT emit it because there is no second automatic self-fix round in which a cause group can recur. **Enforced:** `validators/validate-run.py` `_validate_activity_contract_plan_limits`.
257
274
  - `not-attempted` — the loop never ran because no item qualified.
258
275
  The `no-progress` and `max-rounds-reached` exits are what make the loop terminate; `selfFixMaxRounds` alone is the backstop.
259
- - a `majority-disagree` item with a majority of `needs-user-input` is NOT a self-fix target — it goes straight to the next step's clarification promotion.
276
+ - a `majority-disagree` item with a majority of its deciding `DISAGREE` votes at `needs-user-input` is NOT a self-fix target — after correctness-critical precedence, it goes straight to the next step as `user-decision` rather than generic `noncritical-dissent`. **Enforced:** `validators/validate-run.py` `_expected_approval_classification`.
260
277
  8. For every `majority-disagree` item **that remains after the self-fix loop** (items not resolved by self-fix, or with a `needs-user-input` majority from the start), lead adds a row to `## 1. Clarification Items` with:
261
278
  - new `C-<N>` ID (numbering continues from any existing rows)
262
279
  - `Statement` summarising the disagreement and the worker breakage `<kind>`
263
280
  - `Kind` chosen per the standard policy (usually `decision` for option-level conflicts, `data-point` for path/symbol mismatches)
264
281
  - `Blocks=approval`
265
282
  - the item's `planItems[].clarificationId` set to that `C-<N>` (1:1 link). `validators/validate-run.py` `_validate_plan_body_clarification_matching` recomputes each item's class and fails when a majority-disagree item's `clarificationId` is missing, dangling, or points at a non-`approval` row.
266
- - **A `planner-fixable` item that survives the self-fix loop is NOT promoted to the user by default.** A defect the *planner* could have fixed but did not is still a planner defect; promoting it asks the user to proofread the plan. Once the self-fix budget is exhausted, such an item is recorded as a Working Assumption in `## 5. Missing Information and Risks` — naming the defect, the assumption the implementation will proceed under, and the stop reason — and it stops blocking the gate (it folds into `passed-with-dissent`). It gets **no** `Blocks=approval` row and **no** `clarificationId`.
267
- - **Exception — correctness-critical defects still block.** An item whose `DISAGREE` kinds include `a` (cited path/symbol mismatch), or `f` on a `P-Req-*` item, is promoted per the rules above regardless of `fixability`. These are the defects that make `implementation` produce wrong or unsafe code; `b`/`c`/`e` degrade the plan document, not the resulting code. Rollback ordering (`d`) is advisory — a human runs the rollback — so it is never correctness-critical and never blocks. **Enforced:** `validators/validate-run.py` `_is_dissent_downgraded` (fold) + `_is_correctness_critical` (the exception).
283
+ - set `approvalContext.classification` to `user-decision` for a majority `needs-user-input` item, `correctness-critical` for `DISAGREE(a)`, `DISAGREE(f)` on `P-Req-*`, or an independent Requirement Coverage blocker, and `noncritical-dissent` for another surviving majority disagreement.
284
+ - populate `approvalContext.planItemIds`, `activityIds`, `unblockCondition`, and `recommendedDisposition`. Every option carries a `disposition`: `select` only for `user-decision`, `accept-risk` only for `noncritical-dissent`, and `request-revision` / `reject` for any classification. `correctness-critical` never offers or records `accept-risk`. **Enforced:** `validators/validate-run.py` `_validate_approval_context`.
285
+ - **Self-fix exhaustion is not risk acceptance.** A `noncritical-dissent` item remains blocking until the user explicitly selects `accept-risk`. Record the user's non-empty original text and the `user-decision-required` / `user-decision-evaluated` activity references in `approvalContext.resolution`; only then does `validators/validate-run.py` `_resolved_noncritical_dissent_ids` let `_is_dissent_downgraded` fold it into `passed-with-dissent`.
286
+ - **Correctness-critical defects cannot be waived.** After the user-directed correction, targeted re-verification of every linked item MUST record only `AGREE` or an acceptable `SUPPLEMENT`, and any independent Requirement Coverage blocker MUST be removed before the row becomes `resolved`. A `DISAGREE` or `verification-error` returns it to `open`. **Enforced:** `validators/validate-run.py` `_validate_correctness_resolution`.
268
287
  - When a correctness-critical `planner-fixable` item is promoted, its `Statement` MUST state "planner self-fix attempted but unresolved" and name the stop reason. `validators/validate-run.py` `_validate_self_fix_before_clarification` fails when a planner-fixable majority item is promoted while the budget is not exhausted — it requires `selfFixRoundsApplied >= 1` **and** `selfFixStopReason` in `{no-progress, max-rounds-reached}`, so neither `all-resolved` nor `not-attempted` can excuse a promotion.
288
+ - Approval state transitions are fixed:
289
+ - `open → answered` when the raw user response is recorded
290
+ - `answered → resolved` only after the selected disposition is applied and its checks pass
291
+ - `answered → open` when application or checking fails
292
+ - `open → obsolete` only when a plan change removes the question
293
+ `open` and `answered` continue to block approval; only `resolved` and `obsolete` are non-blocking. A user-directed correction does not consume the automatic self-fix limit, and a failed check does not restart the automatic loop.
294
+ - A terminal row may preserve its original dissent classification only from audited history. Keep superseded votes in `state/plan-body-verification-implementation-planning-<seq>.json`; `validators/validate-run.py` `_historical_plan_item_evidence` recomputes criticality from the recorded `DISAGREE(a|f)` tokens and does not trust the row's classification alone. Every referenced `user-decision-required` / `user-decision-evaluated` activity must cite exactly that row's `C-NNN` and exactly the linked `approvalContext.planItemIds` set. The evaluated activity occurs after every required activity, has `outcome: resolved`, records at least one command whose every `exitCode` is `0`, points `resultPath` at the matching plan-body state artifact, and cites exactly one `plan-body-verification:round-N` evidence token. Round `N` is later than the recorded blocking round; its state votes are all `AGREE` or `SUPPLEMENT`, and they exactly match the final-report verdicts. Its immutable `completedAt` is later than every referenced required activity's canonical event timestamp and no later than every referenced evaluated activity's canonical event timestamp, so an older successful round cannot be relabelled as the response check. This user-response round is recorded in `roundHistory[]` but does not increment `selfFixRoundsApplied` or create another automatic `verification-round-completed` activity. Only the exact round token in `resolution.checkRefs` of a resolved `correctness-critical` row receives that exclusion; the referenced resolved `user-decision-evaluated` activity must match the row's exact `C-NNN` and plan-item set. **Enforced:** `validators/validate-run.py` `_validate_approval_activity_refs`, `_validate_correctness_resolution`, and `_validate_target_round_causality`, plus `validators/validate_session_conformance.py` `_resolved_correctness_reverification_rounds` and `_check_activity_round_counts`. When an independent coverage-only blocker is corrected, keep the `C-NNN` in the now non-blocking Requirement Coverage row's `decisionRefs` and in the matching state-sidecar plan item's `clarificationId`; that item must have no historical blocking dissent and must participate in a round whose `gateBlockedBy` contains `coverage-gap`. A run-wide `coverage-gap` without this item-level `C-NNN` link cannot classify another row. `obsolete` is valid only after current evidence shows that the question or blocker disappeared, or the linked item is historical and removed; a current linked item remains active even when audited history preserves an older classification.
269
295
  9. Approval lives in the report's YAML frontmatter `approved:` field — there is no in-body marker line. The user may flip it to `true` only when the Gate result is `passed` or `passed-with-dissent`. **Enforced:** run-prep (`scripts/okstra_ctl/run.py` `_validate_approved_plan`) fail-closes an `approved: true` plan whose data.json carries a blocking `gateResult` or an open/answered `Blocks: approval` clarification row, and `validators/validate-run.py` `_validate_plan_body_gate_recompute` rejects a declared `gateResult` healthier than the recorded votes.
270
296
 
271
297
  ## `plan-body-verification-<task-type>-<seq>.json` schema
@@ -329,6 +355,7 @@ The per-round structures mirror the finding-convergence state artifact ([converg
329
355
  "roundHistory": [
330
356
  {
331
357
  "round": 1,
358
+ "completedAt": "2026-08-15T01:20:00Z",
332
359
  "gateResult": "blocked-by-disagreement",
333
360
  "gateBlockedBy": ["majority-disagree"],
334
361
  "dispatches": [
@@ -337,6 +364,7 @@ The per-round structures mirror the finding-convergence state artifact ([converg
337
364
  },
338
365
  {
339
366
  "round": 2,
367
+ "completedAt": "2026-08-15T01:35:00Z",
340
368
  "gateResult": "passed-with-dissent",
341
369
  "gateBlockedBy": [],
342
370
  "dispatches": [
@@ -349,7 +377,7 @@ The per-round structures mirror the finding-convergence state artifact ([converg
349
377
 
350
378
  > Abbreviated example: a one-round run has a single `roundHistory[]` entry and a single `rounds[]` entry per item. `P-Opt-1` above is not re-verified in round 2 because the round is focused on the corrected items and the ones the rewrite touched (step 7) — an item may legitimately carry fewer `rounds[]` entries than `roundHistory[]` has rounds, but every round in `roundHistory[]` must appear on at least one item.
351
379
 
352
- `roundHistory[].gateResult` / `gateBlockedBy` are that round's own gate resolution (§"Round protocol" step 5), not the run's final one — the final value lives in data.json. `dispatches[].terminalStatus` mirrors finding convergence (`completed | timeout | error | not-run`). A wrapper-recorded `cli-failure` is a run-error-log event, not a terminal status — record that dispatch's `terminalStatus` as `error`.
380
+ `roundHistory[].gateResult` / `gateBlockedBy` are that round's own gate resolution (§"Round protocol" step 5), not the run's final one — the final value lives in data.json. `roundHistory[].completedAt` is the immutable ISO 8601 UTC timestamp recorded once after the round passes `okstra plan-verify`; it is not copied from a later activity and is never revised. `dispatches[].terminalStatus` mirrors finding convergence (`completed | timeout | error | not-run`). A wrapper-recorded `cli-failure` is a run-error-log event, not a terminal status — record that dispatch's `terminalStatus` as `error`.
353
381
 
354
382
  `planItems[].rounds[].classification` enum: `full-consensus | partial-consensus | dissent-isolated | majority-disagree | needs-reverify | contested`. `needs-reverify` is the peer-error shape from §"Round protocol" step 4 (a single-vote-blocking kind with fewer than 2 participating non-error votes) — it survives into the state file when the round budget runs out before the re-dispatch resolves it, and `_recompute_plan_body_gate` folds it into `passed-with-dissent`. `contested` only appears when `maxRounds > 1`; at default `maxRounds=1` any otherwise-unresolved item folds into `partial-consensus` per the round protocol above.
355
383