okstra 0.171.0 → 0.173.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (106) hide show
  1. package/README.md +8 -6
  2. package/docs/architecture/storage-model.md +11 -0
  3. package/docs/architecture.md +29 -14
  4. package/docs/cli.md +40 -7
  5. package/docs/for-ai/skills/okstra-user-response.md +2 -2
  6. package/docs/performance-improvement-plan-v2.md +6 -5
  7. package/docs/project-structure-overview.md +24 -14
  8. package/docs/task-process/README.md +5 -3
  9. package/docs/task-process/error-analysis.md +2 -2
  10. package/docs/task-process/final-verification.md +2 -2
  11. package/docs/task-process/implementation-option-selection.md +70 -0
  12. package/docs/task-process/implementation-planning.md +23 -15
  13. package/docs/task-process/requirements-discovery.md +2 -2
  14. package/package.json +1 -1
  15. package/runtime/BUILD.json +2 -2
  16. package/runtime/agents/workers/report-writer-worker.md +30 -6
  17. package/runtime/bin/lib/okstra/cli.sh +5 -1
  18. package/runtime/bin/lib/okstra/globals.sh +1 -0
  19. package/runtime/bin/lib/okstra/usage.sh +3 -0
  20. package/runtime/bin/okstra.sh +2 -0
  21. package/runtime/prompts/duties/direction-selection-worker.md +44 -0
  22. package/runtime/prompts/duties/planning-worker.md +12 -4
  23. package/runtime/prompts/launch.template.md +4 -0
  24. package/runtime/prompts/lead/adapters/cmux.md +1 -1
  25. package/runtime/prompts/lead/context-loader.md +1 -1
  26. package/runtime/prompts/lead/convergence.md +5 -5
  27. package/runtime/prompts/lead/okstra-lead-contract.md +42 -17
  28. package/runtime/prompts/lead/plan-body-verification.md +42 -14
  29. package/runtime/prompts/lead/report-writer.md +38 -15
  30. package/runtime/prompts/lead/team-contract.md +2 -0
  31. package/runtime/prompts/profiles/_clarification-recommendation.md +3 -1
  32. package/runtime/prompts/profiles/_common-contract.md +3 -2
  33. package/runtime/prompts/profiles/_implementation-deliverable.md +2 -2
  34. package/runtime/prompts/profiles/error-analysis.md +3 -3
  35. package/runtime/prompts/profiles/final-verification.md +3 -3
  36. package/runtime/prompts/profiles/forbidden-actions.json +7 -0
  37. package/runtime/prompts/profiles/implementation-option-selection.md +35 -0
  38. package/runtime/prompts/profiles/implementation-planning.md +56 -37
  39. package/runtime/prompts/profiles/implementation.md +2 -1
  40. package/runtime/prompts/profiles/improvement-discovery.md +1 -1
  41. package/runtime/prompts/profiles/requirements-discovery.md +3 -3
  42. package/runtime/prompts/wizard/prompts.ko.json +9 -1
  43. package/runtime/python/okstra_ctl/adapters/hosts/antigravity/relay.md +1 -1
  44. package/runtime/python/okstra_ctl/adapters/hosts/claude-code/relay.md +1 -1
  45. package/runtime/python/okstra_ctl/adapters/hosts/codex/relay.md +1 -1
  46. package/runtime/python/okstra_ctl/adapters/hosts/external/relay.md +1 -1
  47. package/runtime/python/okstra_ctl/adapters/hosts/grok/relay.md +1 -1
  48. package/runtime/python/okstra_ctl/adapters/hosts/kimi/relay.md +1 -1
  49. package/runtime/python/okstra_ctl/agent_activity.py +306 -0
  50. package/runtime/python/okstra_ctl/agent_invocation.py +1 -0
  51. package/runtime/python/okstra_ctl/analysis_packet.py +6 -0
  52. package/runtime/python/okstra_ctl/clarification_items.py +37 -20
  53. package/runtime/python/okstra_ctl/exact_coverage.py +128 -0
  54. package/runtime/python/okstra_ctl/fix_cycles.py +3 -1
  55. package/runtime/python/okstra_ctl/implementation_direction.py +836 -0
  56. package/runtime/python/okstra_ctl/implementation_options.py +479 -0
  57. package/runtime/python/okstra_ctl/lead_events.py +47 -4
  58. package/runtime/python/okstra_ctl/plan_items.py +51 -3
  59. package/runtime/python/okstra_ctl/render.py +12 -3
  60. package/runtime/python/okstra_ctl/render_final_report.py +1 -0
  61. package/runtime/python/okstra_ctl/report_contract.py +45 -13
  62. package/runtime/python/okstra_ctl/report_finalize.py +51 -14
  63. package/runtime/python/okstra_ctl/report_html/common.py +5 -3
  64. package/runtime/python/okstra_ctl/report_html/render.py +4 -2
  65. package/runtime/python/okstra_ctl/report_html/router.py +4 -0
  66. package/runtime/python/okstra_ctl/report_html/view_models/implementation_option_selection.py +32 -0
  67. package/runtime/python/okstra_ctl/report_html/view_models/implementation_planning.py +42 -11
  68. package/runtime/python/okstra_ctl/report_translation.py +14 -0
  69. package/runtime/python/okstra_ctl/report_views.py +148 -12
  70. package/runtime/python/okstra_ctl/run.py +350 -2
  71. package/runtime/python/okstra_ctl/scope_provenance.py +15 -9
  72. package/runtime/python/okstra_ctl/user_response.py +75 -0
  73. package/runtime/python/okstra_ctl/wizard.py +144 -0
  74. package/runtime/python/okstra_ctl/worker_audit_ledger.py +150 -0
  75. package/runtime/python/okstra_ctl/worker_prompt_policy.py +2 -0
  76. package/runtime/python/okstra_ctl/workflow.py +29 -7
  77. package/runtime/schemas/final-report-v2.0.schema.json +1623 -143
  78. package/runtime/skills/okstra-user-response/SKILL.md +2 -2
  79. package/runtime/templates/reports/final-report-v2.template.md +12 -0
  80. package/runtime/templates/reports/final-verification-input.template.md +1 -1
  81. package/runtime/templates/reports/html/assets/base.css +7 -0
  82. package/runtime/templates/reports/html/base.template.html +3 -2
  83. package/runtime/templates/reports/html/i18n/en.json +27 -2
  84. package/runtime/templates/reports/html/i18n/ko.json +27 -2
  85. package/runtime/templates/reports/html/macros/forms.html +42 -4
  86. package/runtime/templates/reports/html/tasks/implementation-option-selection.template.html +49 -0
  87. package/runtime/templates/reports/html/tasks/implementation-planning.template.html +61 -2
  88. package/runtime/templates/reports/i18n/en.json +17 -0
  89. package/runtime/templates/reports/implementation-input.template.md +4 -2
  90. package/runtime/templates/reports/implementation-planning-input.template.md +18 -4
  91. package/runtime/templates/reports/improvement-discovery-input.template.md +1 -1
  92. package/runtime/templates/reports/md/tasks/implementation-option-selection.template.md +13 -0
  93. package/runtime/templates/reports/md/tasks/implementation-planning.template.md +17 -0
  94. package/runtime/templates/reports/report.js +137 -21
  95. package/runtime/templates/reports/task-brief.template.md +9 -3
  96. package/runtime/templates/reports/user-response.template.md +28 -5
  97. package/runtime/templates/worker-prompt-preamble.md +16 -0
  98. package/runtime/validators/validate-implementation-plan-stages.py +106 -1
  99. package/runtime/validators/validate-report-views.py +2 -2
  100. package/runtime/validators/validate-run.py +1124 -54
  101. package/runtime/validators/validate_improvement_report.py +5 -1
  102. package/runtime/validators/validate_session_conformance.py +523 -35
  103. package/src/cli-registry.mjs +7 -0
  104. package/src/commands/execute/codex-run.mjs +1 -0
  105. package/src/commands/execute/render-bundle.mjs +1 -0
  106. package/src/commands/report/agent-activity.mjs +21 -0
@@ -8,7 +8,15 @@ The JSON SSOT path is `runs/<task-type>/reports/final-report-<task-type>-<seq>.d
8
8
 
9
9
  New bundles use `schemas/final-report-v2.0.schema.json`. The Markdown keeps verdict, routing, evidence, one structured task deliverable, and audit data for the next agent. The HTML uses `humanSummary`, task `userNarrative`, and structured facts for the user. Raw worker discussion, convergence mechanics, and usage belong to audit structures and never to the HTML human main body.
10
10
 
11
- Two `frontmatter` approval fields are always emitted with their unset default — never pre-fill them: `frontmatter.approved` is emitted as `false`, and `frontmatter.implementationOption` is emitted as an empty string `""`. The user later flips `approved` to `true` (via `--approve` or manual edit) and fills `implementationOption` with the chosen Option Candidate name (via `--implementation-option <name>` or manual edit) to authorise and scope the next `implementation` run.
11
+ ### Implementation-planning frontmatter contract
12
+
13
+ #### Selected-direction
14
+
15
+ Emit `frontmatter.approved` as `false` and copy `implementationPlanning.selectedDirectionRef.snapshotPath` into `frontmatter.selectedDirectionRef`. You MUST omit `frontmatter.implementationOption`; the direction was selected upstream and cannot be selected again in planning. `schemas/final-report-v2.0.schema.json` enforces the required selected-direction reference and rejects an `implementationOption` property for this branch.
16
+
17
+ #### Legacy candidate-comparison
18
+
19
+ Emit `frontmatter.approved` as `false` and `frontmatter.implementationOption` as the empty string `""`. The user later flips `approved` to `true` and fills `implementationOption` with the chosen Option Candidate name to authorise and scope the next `implementation` run. Every other report type follows the same empty `implementationOption` default; the schema's non-selected-direction branch requires that field and rejects a selected-direction reference.
12
20
 
13
21
  **As the report-writer worker:** YOU write the data.json and invoke the renderer; the files on disk are the canonical record, so do not return either artifact inline.
14
22
 
@@ -85,11 +93,11 @@ For an implementation-planning run, the Report writer worker owns the Phase 6 de
85
93
 
86
94
  ### Before `report-finalize`: the translation sidecar (BLOCKING order)
87
95
 
88
- `report-finalize` step `render-views` overlays the translation sidecar, so a non-English run must produce that sidecar **before** the command runs. That leaves exactly one correct order, and it is not the intuitive one:
96
+ The finalization renderer overlays the translation sidecar, so a non-English run must produce that sidecar before finalization. Use this fixed order:
89
97
 
90
- 1. **Verify the data.json is English first.** Run `okstra report-translate check-source <data.json>`. Do this even when **Report Language** is `en` — it is the cheapest gate in the phase and it protects every step after it.
91
- 2. **Only when it passes and Report Language is not `en`**, dispatch the translator worker, which writes `final-report-<task-type>-<seq>.i18n.<lang>.json`.
92
- 3. Then run `report-finalize`.
98
+ 1. **For a non-English report only**, run `okstra report-finalize ... --only project-activity --only check-source`. The shared finalizer projects canonical activity before checking the English source. A historical manifest without `activityContractVersion: 1` leaves data.json unchanged.
99
+ 2. **Only when that check passes**, dispatch the translator worker, which writes `final-report-<task-type>-<seq>.i18n.<lang>.json`.
100
+ 3. Run the full `report-finalize` command below. English reports start here; the finalizer repeats the idempotent projection and source check before every downstream step.
93
101
 
94
102
  For step 2, write translator-only task instructions and run `okstra
95
103
  agent-prompt materialize` with `--audience translator`, `--assignment-ref
@@ -116,25 +124,26 @@ okstra report-finalize \
116
124
  --report <runDirectoryPath>/reports/final-report-<task-type>-<seq>.md
117
125
  ```
118
126
 
119
- Do NOT run the five steps below by hand. Hand-running them is the recurring root cause of reports shipping with `--` token cells, a missing html sibling, Section 3 missing follow-up entries, or Section 4 rows never spawning — the order is load-bearing and a skipped step surfaces only later, as a validator `contract-violated`. Every step is idempotent, so after fixing a reported failure just re-run the same command.
127
+ Do NOT run the six steps below by hand. Hand-running them is the recurring root cause of reports shipping with stale activity, `--` token cells, a missing html sibling, Section 3 missing follow-up entries, or Section 4 rows never spawning — the order is load-bearing and a skipped step surfaces only later, as a validator `contract-violated`. Every step is idempotent, so after fixing a reported failure just re-run the same command.
120
128
 
121
129
  The steps it executes, in this contractual order, and the contract each one carries:
122
130
 
123
- 1. **`check-source` — verify the data.json is English.** The same gate as the pre-translator check above, run again here because everything after it derives from the data.json: rendering a Korean SSOT into English chrome, spawning follow-ups from it, and validating it all succeed on a record the next phase cannot read. A failure here means the report-writer authored in the reader's language; re-dispatch it with the English rule rather than editing the data.json by hand.
124
- 2. **`token-usage` — collect usage.** Aggregates `leadUsage` / `workers[].usage` / `usageSummary` into team-state, populates `tokenUsage` and the execution-status usage fields in data.json, and re-invokes the renderer so the markdown carries real numbers.
131
+ 1. **`project-activity` — project canonical activity.** Replaces only `agentActivity[]` from this run's canonical events before translation source extraction. A legacy manifest without activity contract v1 is a byte-preserving no-op. Conformance compares IDs, order, and every core field against the canonical events.
132
+ 2. **`check-source` — verify the data.json is English.** The same gate as the pre-translator check above, run again here because everything after it derives from the data.json: rendering a Korean SSOT into English chrome, spawning follow-ups from it, and validating it all succeed on a record the next phase cannot read. A failure here means the report-writer authored in the reader's language; re-dispatch it with the English rule rather than editing the data.json by hand.
133
+ 3. **`token-usage` — collect usage.** Aggregates `leadUsage` / `workers[].usage` / `usageSummary` into team-state, populates `tokenUsage` and the execution-status usage fields in data.json, and re-invokes the renderer so the markdown carries real numbers.
125
134
 
126
135
  The data.json paths populated: `tokenUsage.lead.{totalTokens,billableTokens,costUsd}`, the `worker` / `grand` rows, `tokenUsage.cli.costUsd`, and each `executionStatus[].{totalTokens,billableTokens,costUsd,durationMs,cliTotalTokens,cliCostUsd}` for rows whose role matches a team-state worker. The data.json MUST already exist (Phase 6 output).
127
136
 
128
137
  For implementation-planning, this Phase 7 canonical render calls `materialize_design_prep_requests()` after token substitution and creates deterministic request files only for `provisional` / `blocked` items. Later answers are append-only user-input sidecars; request generation and user input never rewrite the assessment fields, so the source report remains immutable as the design-input snapshot after this render. `validators/validate-run.py` `_validate_design_prep_requests` enforces request existence, canonical path, content, and assessment fingerprint.
129
- 3. **`render-views` — render the human report artifact.** Runs against the substituted v2 data.json and its Markdown sibling.
138
+ 4. **`render-views` — render the human report artifact.** Runs against the substituted v2 data.json and its Markdown sibling.
130
139
 
131
140
  Output (idempotent — re-running overwrites):
132
141
  - `runs/<task-type>/reports/final-report-<task-type>-<seq>.html` — single-file self-contained human view, always generated for schema v2 from the dedicated template registered for that task type. Clarification rows with `Status` ∈ {`open`, `answered`} embed response controls and export a `user-response-<task-type>-<seq>.md` sidecar. The original data and Markdown artifacts are never mutated by user input.
133
- - the implementation-planning report renders a **Plan Approval** section at the end of the body (implementation-option `<select>` + an approval checkbox) — disabled while any §1 `Blocks: approval` row is unresolved. Checking approval and exporting embeds a `## APPROVAL` block in the sidecar body, and the implementation-start wizard's approve-confirm step detects it and, after user confirmation, applies it through the existing `--approve` / `--implementation-option` path.
142
+ - the implementation-planning report renders a **Plan Approval** section at the end of the body — an implementation-option `<select>` plus approval checkbox for legacy candidate-comparison, and an approval checkbox only for selected-direction plans. It stays disabled while any §1 `Blocks: approval` row is unresolved.
134
143
  - Schema-v1 and quick compatibility reports retain the legacy conditional HTML path; this does not change the schema-v2 always-generated contract.
135
144
 
136
145
  It runs after usage collection so token placeholders are substituted in any rendered html, and before routing persistence so the html artifact, when generated, exists for the validator step that checks it. It also overlays the translation sidecar, which is why a non-English run must dispatch the translator before this command — see the ordering rule above.
137
- 4. **`spawn-followups` — routing and follow-up persistence.** Turns the report's `## 4. Follow-up Tasks` rows into `tasks/<task-group>/<new-task-id>/` stubs.
146
+ 5. **`spawn-followups` — routing and follow-up persistence.** Turns the report's `## 4. Follow-up Tasks` rows into `tasks/<task-group>/<new-task-id>/` stubs.
138
147
 
139
148
  Behaviour contract:
140
149
  - Idempotent: rows whose target dir exists are reported as `existing` and skipped. Reruns of the same parent task are safe.
@@ -149,7 +158,7 @@ The steps it executes, in this contractual order, and the contract each one carr
149
158
  ```
150
159
 
151
160
  The status file is written after routing and follow-up persistence completes.
152
- 5. **`validate-run` — validate the finished run.** Checks the completed artifact set, including the report-views contract that catches a missing or stale html sibling. A failure here names the specific contract; fix it and re-run `okstra report-finalize`.
161
+ 6. **`validate-run` — validate the finished run.** Checks the completed artifact set, including exact canonical-event-to-`agentActivity[]` conformance and the report-views contract that catches a missing or stale html sibling. A failure here names the specific contract; fix it and re-run `okstra report-finalize`.
153
162
 
154
163
  After `okstra report-finalize` reports `"ok": true`, **execute the run-scoped cleanup gate.** Call `shutdown_workers` only after that success, all persistence work, and explicit user approval under [okstra-lead-contract](./okstra-lead-contract.md) "Run-scoped worker-resource lifecycle". If the user keeps resources, leave the selected adapter's resources intact and surface its manual cleanup guidance.
155
164
 
@@ -292,7 +301,7 @@ When the run's `task-type` is `final-verification`, the report's `## 7. Final Ve
292
301
  |---|--------------------|---------|
293
302
  | 1 | `accepted` | All acceptance criteria pass; `release-handoff` may proceed. |
294
303
  | 2 | `conditional-accept` | Acceptance passes with caveats; user must resolve listed conditions before `release-handoff`. |
295
- | 3 | `blocked` | Acceptance failed; routing returns to `error-analysis` or `implementation-planning`. |
304
+ | 3 | `blocked` | Acceptance failed; routing returns to `error-analysis`, `implementation-option-selection`, or `implementation-planning` according to whether the cause, direction, or detailed plan failed. |
296
305
 
297
306
  For every other task-type, set the `Verdict Token` cell to `not-applicable`. Do NOT omit the row — the template renders it for all task-types and downstream tooling expects the field to exist.
298
307
 
@@ -346,12 +355,26 @@ Every field MUST anchor its claim with at least one evidence reference — a `pa
346
355
  1. **Clarification Items** — single unified `C-*` table; column schema (4 columns with the short fields stacked in one record-meta cell), ID convention, and rerun behaviour are owned by `_common-contract.md §Clarification request policy` (SSOT). The deprecated `5.5.9 Open Questions` / `1.1 Additional Material Request` / `1.2 User Confirmation Questions` sub-sections are removed; the validator fails reports that reintroduce them.
347
356
  - **Open `Blocks=approval` rows carry `origin` and `userConfirmation`** (same SSOT). Lead's dispatch prompt MUST state, per intended blocker, which `origin` applies and what Lead did about it — the writer cannot observe either. When Lead instructed the writer to raise an item rather than decide it, that row's `origin` is `lead-directed` no matter how the workers subsequently voted on it: an instruction returning as a consensus is not a finding. Before writing such an instruction, run the confirmation sequence in [okstra-lead-contract](./okstra-lead-contract.md) "User confirmation before an approval blocker" — asking first is usually cheaper than the row.
348
357
  2. **Evidence and Detailed Analysis** — primary evidence rows (file path, line, snippet); secondary evidence / alternate interpretations. If `reference-expectations.md` lists explicit expected values, record match/gap per row.
349
- - **Error-analysis diagnosis and routing.** When `header.taskType` is `error-analysis`, populate the required `errorAnalysis` object. Copy `errorAnalysis.symptomVerbatim` byte-for-byte from the symptom stated in the brief's `Source Material`; do not paraphrase it. Every `causeCandidates[]` row includes the full `supportingEvidence`, `falsifyingEvidenceChecked`, `confidence`, and `disproveWith` fields. When a candidate is a step in a propagation chain rather than a competing explanation — the analysis calls it a downstream step, a second stage, or a consequence of another candidate — set its `downstreamOf` to the ids of the candidates immediately upstream of it; leave the field absent for a candidate that stands on its own. Every id listed MUST be another candidate in the same report, no row may name itself, and the links MUST NOT form a cycle; `validators/validate-run.py::_validate_cause_chain` rejects all three. This is the only place the chain is machine-readable — prose calling a candidate "the second step of the chain" while `downstreamOf` is absent leaves the report's figure claiming the candidates are alternatives. Route `errorAnalysis.routing.nextTaskType=implementation-planning` with `direction=begin-planning`, or route `errorAnalysis.routing.nextTaskType=error-analysis` with `direction=continue-investigation`; no other pairing is valid. `verdictCard.nextStep`, `finalVerdict.nextStep`, the first `recommendedNextSteps` action and command, and the unique `followUpTasks` row whose `origin` is `phase-continuation` MUST all point to the same `errorAnalysis.routing.nextTaskType` target. The schema enforces only the presence of a `phase-continuation` row. Phase validation MUST enforce exact target agreement and uniqueness through `validators/validate-run.py::_validate_error_analysis_consistency`; until that check is implemented and executed, those semantics are contract requirements rather than enforced guarantees.
358
+ - **Error-analysis diagnosis and routing.** When `header.taskType` is `error-analysis`, populate the required `errorAnalysis` object. Copy `errorAnalysis.symptomVerbatim` byte-for-byte from the symptom stated in the brief's `Source Material`; do not paraphrase it. Every `causeCandidates[]` row includes the full `supportingEvidence`, `falsifyingEvidenceChecked`, `confidence`, and `disproveWith` fields. When a candidate is a step in a propagation chain rather than a competing explanation — the analysis calls it a downstream step, a second stage, or a consequence of another candidate — set its `downstreamOf` to the ids of the candidates immediately upstream of it; leave the field absent for a candidate that stands on its own. Every id listed MUST be another candidate in the same report, no row may name itself, and the links MUST NOT form a cycle; `validators/validate-run.py::_validate_cause_chain` rejects all three. This is the only place the chain is machine-readable — prose calling a candidate "the second step of the chain" while `downstreamOf` is absent leaves the report's figure claiming the candidates are alternatives. Route `errorAnalysis.routing.nextTaskType=implementation-option-selection` with `direction=begin-option-selection`, or route `errorAnalysis.routing.nextTaskType=error-analysis` with `direction=continue-investigation`; no other pairing is valid. `verdictCard.nextStep`, `finalVerdict.nextStep`, the first `recommendedNextSteps` action and command, and the unique `followUpTasks` row whose `origin` is `phase-continuation` MUST all point to the same `errorAnalysis.routing.nextTaskType` target. The schema enforces only the presence of a `phase-continuation` row; `validators/validate-run.py::_validate_error_analysis_consistency` enforces exact target agreement and uniqueness.
359
+ - **Implementation-option-selection comparison.** When `header.taskType` is `implementation-option-selection`, populate `implementationOptionSelection` from the converged direction-selection findings. Preserve every merged or rejected raw candidate in `candidateAudit`, and put at most three selectable candidates in `rankedOptions`. Each displayed candidate carries its requirement coverage, scope commitments, criterion scores, feasibility votes, safety blockers, unresolved feasibility facts, planning invariants, and exact coverage summary. In each displayed candidate, `expectedChangeAreas` names direction-level change surfaces, never exact file paths or an exact file list. `expectedVerification` names direction-level verification signals, never a stage list or executable test commands. `schemas/final-report-v2.0.schema.json` enforces the displayed-summary constants and the three-option cap; semantic recalculation belongs to `validators/validate-run.py`.
360
+ - **Implementation-planning direction branch.** When `implementationPlanning.planningContract == "selected-direction"`, read `selectedDirectionRef` and the snapshot before authoring. Materialize the snapshot into `directionRealization`, stages, validation, rollback, and bidirectional original-requirement links. Author exactly one `P-Dir-1`; its payload is the complete `directionRealization`. Its verification covers the core mechanism, architecture boundaries, planning invariants, and any hidden direction change against `selectedDirectionRef`. Do not author Option Candidates, candidate scores, a Recommended Option, or user candidate-selection fields. When current evidence requires changing the direction, author `outcome: "direction-invalidated"` and omit the execution plan. Legacy candidate-comparison reruns retain `P-Opt-*`, Option Candidates, trade-off, and Recommended Option semantics.
361
+
362
+ ```json
363
+ {
364
+ "candidateDetailBoundary": {
365
+ "expectedChangeAreas": "direction-level-only",
366
+ "expectedVerification": "direction-level-signals-only",
367
+ "forbidden": ["exact-file-lists", "stage-lists", "test-commands"]
368
+ }
369
+ }
370
+ ```
371
+
372
+ - **Implementation-option-selection is non-terminal.** Its `followUpTasks` includes a `phase-continuation` row with `autoSpawn: "no"` and `priority: "P0"`; the schema's non-terminal conditional enforces row presence.
350
373
  3. **Recommended Next Steps** — prioritized actions. After Phase 7's follow-up spawner runs, append a row per newly created task-key (see "Phase 6 → Phase 7 execution sequence" above). **Approval-gate consistency:** when §1 carries any `Blocks: approval` row with `Status` ∈ {open, answered}, the Verdict Card `Next Step` and the first recommended step MUST point to the clarification rerun (`resume-clarification` of the SAME task-type) — never to "flip frontmatter `approved: true` → jump straight to `implementation`". Run-prep enforces this gate (`run.py _validate_approved_plan` fail-closes on those rows and on a blocking data.json `gateResult`), so a direct-implementation next-step is an instruction the reader cannot actually follow. **Cross-project pointer rule:** for cross-project dependencies (another repo / a different top-level deployment module / a published package), `crossProjectDependencies` (§5.4 Cross-Project Dependencies) is authoritative — do NOT duplicate that substance (prerequisite work / verification signals / handoff) into `recommendedNextSteps`; put only a one-line pointer to that section (no double-recording).
351
374
  4. **Follow-up Tasks** — auto-spawn-eligible table. Each row drives `okstra-spawn-followups.py`; see template §4 for the row schema.
352
375
  5. **Missing Information and Risks** — uncertain / "I don't know" items. `implementation-planning` adds §5.5 (see heading contract below); `release-handoff` adds §5.6.
353
376
  6. **Cross Verification Results** — 4 categories (Full / Partial / Contested / Worker-Unique) when convergence is enabled, per `convergence`. Prepend the Round History sub-table (columns: `Round | inputQueueSize | resolvedCount | carriedForwardCount | dispatches | skippedWorkers`) plus a `round2SkippedReason: <value>` note, pulled verbatim from `convergence-<task-type>-<seq>.json`. Empty contested list renders as `- No items lacking consensus.`. Convergence-disabled runs use the legacy Consensus/Differences format and omit the round table.
354
- 7. **Final Verdict** — `Direction` ∈ `continue-investigation` / `begin-planning` / `begin-implementation` / `approve` / `reject` / `hold`. **Verdict Token** is `not-applicable` for every task-type except `final-verification` — see "Final-verification verdict token contract" below for that case.
377
+ 7. **Final Verdict** — `Direction` ∈ `continue-investigation` / `begin-option-selection` / `begin-planning` / `begin-implementation` / `approve` / `reject` / `hold`. **Verdict Token** is `not-applicable` for every task-type except `final-verification` — see "Final-verification verdict token contract" below for that case.
355
378
 
356
379
  **§5.10 Fix History (data-presence gated).** When the run-manifest carries a `fixCycleId`, fill the data.json `fixCycle` block (`cycle` / `targetReport` / `symptom` / `runs`). Read the values from the task root's `history/fix-cycles.jsonl`: `cycle` MUST equal `fixCycleId`, `targetReport` / `symptom` come from that cycle's `opened` row, and `runs` lists its attached `run` rows (`taskType` / `runSeq` / `runManifest`). The validator (`validators/validate-run.py` → `_validate_fix_cycle`) fails the run when the block is missing or `fixCycle.cycle` does not match `fixCycleId`. When the run-manifest has no `fixCycleId`, OMIT the `fixCycle` block entirely — the renderer omits §5.10.
357
380
 
@@ -164,6 +164,8 @@ After each worker attempt returns (regardless of role), Lead MUST verify the can
164
164
  - The result file exists but its audit sidecar does not, at `runs/<task-type>/worker-results/<worker>-audit-<task-type>-<seq>.md`. Workers write both in the same step, so a result without a sidecar means the Reading Confirmation block — the only evidence the worker read its inputs — was never produced. `validate-run.py` fails the run on this at Phase 7 either way (`validate_worker_results_audit`); checking it here spends the existing one-retry budget while the role can still be re-dispatched, instead of surfacing hours later when the worker session is gone.
165
165
  - `okstra worker-audit-check --run-dir <runs/<task-type>/> --task-type <t> --seq <n> --worker <id>` exits 2 on a backticked `path:line` citation in the worker's result that has no matching Evidence read row in its audit sidecar. Run it the moment you collect each result. Phase 7 enforces the same rules from the same implementation (`okstra_ctl.worker_audit_ledger`), but by then the worker session is gone and the only remaining moves are editing the result yourself — which destroys the audit chain the ledger exists to provide — or ending the run `contract-violated`. While the session is alive, `SendMessage` to the worker so it corrects its own citation; that costs about a minute against a re-dispatch or a failed run.
166
166
 
167
+ The same audit check parses command evidence from canonical rows such as `- Evidence command: {"command":"npm run check","cwd":"<project-root>","exitCode":0,"outputSummary":"all checks passed"}`. Workers record only commands that produced or verified a conclusion, not exploratory `rg`, `ls`, or file-opening commands. Environment variable values, tokens, credentials, and authorization headers are excluded. A malformed row or potential sensitive material is a contract failure returned by `okstra_ctl.worker_audit_ledger`; send the failure to the worker while its session is still available.
168
+
167
169
  **One-retry policy:**
168
170
 
169
171
  1. On the FIRST result-missing trigger for a given role within a single run, Lead MUST call `redispatch_worker` with the byte-identical prompt — same `**Result Path:**`, same `**Prompt History Path:**`, same model assignment, and same adapter-native assignment identity. The redispatch counts as a second attempt against the existing role slot; do NOT create a new role-id, do NOT change the result file path, do NOT switch to a different model as a "workaround".
@@ -1,10 +1,12 @@
1
- - every `Kind=decision` clarification row carries its choices in `options[]`, never as prose inside `expectedForm`. Each option is an object with six fields:
1
+ - every `Kind=decision` clarification row carries its choices in `options[]`, never as prose inside `expectedForm`. Each option is an object with seven fields:
2
2
  - `role` — `recommended` for the single best answer, `alternative` for the rest. Exactly one option per row is `recommended`.
3
3
  - `answer` — the choice itself, phrased so the user can pick it as-is. Keep it to a short phrase (roughly 120 characters); the reasoning and the consequences have their own fields below.
4
4
  - `rationale` — one sentence on why this option is on the board.
5
5
  - `scopeImpact` — tokens drawn from `{in-repo, cross-repo, new-schema, deferrable}`. Exactly one of `in-repo` / `cross-repo`, which answer the same question and are mutually exclusive; `new-schema` and `deferrable` are optional additions.
6
6
  - `addedWork` — one sentence naming the work this choice creates that the other choices do not. Name the work, not a cost adjective.
7
7
  - `directionChange` — one sentence naming what this choice reverses: an approved plan item, a recorded decision, an earlier answer. When it reverses nothing, say so.
8
+ - `disposition` — the effect of selecting the option. Use `select` for `user-decision`, `accept-risk` for `noncritical-dissent`, and `request-revision` or `reject` when the option sends the plan back. `correctness-critical` never offers `accept-risk`.
9
+ - an approval-blocking row carries `approvalContext`. Its `classification` is one of `user-decision`, `noncritical-dissent`, or `correctness-critical`; `planItemIds`, `activityIds`, `unblockCondition`, and `recommendedDisposition` are required. A completed decision records `resolution` with `disposition`, non-empty `userText`, and verification `checkRefs`.
8
10
  - the three impact fields answer three different questions — how far the change reaches, what new work it creates, and what it overturns. Someone choosing between options needs all three, so never fold them into one sentence: whichever axis is easiest to write would silently stand in for the other two.
9
11
  - a row that omits `options[]`, offers fewer than two, or marks zero or two options as `recommended` is incomplete and must be completed before the report is finalised.
10
12
  - `expectedForm` states only the *shape* of the answer — one of the options, a file path, a number, a date. It never lists the choices again; two sources for one fact leave consumers disagreeing about which is authoritative.
@@ -9,8 +9,9 @@ profile document.
9
9
  - Worker interaction model (shared — read before inferring behaviour from the roster):
10
10
  - the per-profile `Required workers:` block is a **roster**, not a behaviour contract. Each role's interaction mode changes across operating phases of the same run.
11
11
  - **Phase 4 / 5 (independent analysis)**: every analyser in the resolved provider assignment roster produces findings independently and has no access to another worker's output. `report-writer` does not analyse.
12
- - **Phase 5.5 (convergence — peer review by workers)**: workers peer-review each other's findings across up to `effectiveMaxRounds` rounds; the lead mediates but does not vote. See `prompts/lead/convergence.md` for the round protocol (replay of findings, `AGREE` / `DISAGREE` / `SUPPLEMENT` verdicts), queue invariants, and final classification (`full-consensus` / `partial-consensus` / `contested` / `worker-unique`). For `requirements-discovery`, `error-analysis`, `implementation-planning`, `project-analysis`, `feature-analysis`, and `change-impact-analysis` this phase runs in **adversarial mode** (`convergence.adversarial=true`): verifiers try to refute each finding against its cited evidence and the burden of proof sits on the claim — see that skill's §"Adversarial Verification Mode".
12
+ - **Phase 5.5 (convergence — peer review by workers)**: workers peer-review each other's findings across up to `effectiveMaxRounds` rounds; the lead mediates but does not vote. See `prompts/lead/convergence.md` for the round protocol (replay of findings, `AGREE` / `DISAGREE` / `SUPPLEMENT` verdicts), queue invariants, and final classification (`full-consensus` / `partial-consensus` / `contested` / `worker-unique`). For `requirements-discovery`, `error-analysis`, `implementation-option-selection`, `implementation-planning`, `project-analysis`, `feature-analysis`, and `change-impact-analysis` this phase runs in **adversarial mode** (`convergence.adversarial=true`): verifiers try to refute each finding against its cited evidence and the burden of proof sits on the claim — see that skill's §"Adversarial Verification Mode".
13
13
  - Do NOT conclude "no peer review happens" from the roster alone — every profile that lists ≥2 analyser workers runs convergence by default (`convergence.enabled=true` in `task-manifest.json`).
14
+ - For a new `implementation-planning` run, the plan-body sequence is initial verification → one planner self-fix → targeted re-verification → user gate. The initial verification is round 1, the targeted re-verification is round 2, and a second automatic self-fix is a contract violation. A user-directed correction does not consume the automatic self-fix limit, and a verification failure after that correction does not restart the automatic loop.
14
15
  - **provider-unavailable fallback (tolerance).** A worker dispatch can fail to produce a result for two distinct reasons, and both take the same recovery path. (1) **Pane budget:** the dispatch is rejected with `no room for another tmux split` (or an equivalent teammate-pane creation failure). (2) **Sandbox CLI-start failure (non-tmux path):** an external CLI worker wrapper exits non-zero within seconds with empty stdout and its live-log shows `operation not permitted`. In either case the lead spends the one shared retry budget through the assignment's recorded runner. If the provider is still unavailable, record that terminal status and continue only under the convergence quorum rules; never replace it silently with a fixed provider or count a substitute as the original provider's vote. Completed external-CLI worker panes are reclaimed by the selected runtime adapter's resource lifecycle. (This is a prompt instruction, not a code-enforced gate.)
15
16
  - Dual-audience final-report contract (shared):
16
17
  - data.json is the sole authored report artifact. AI handoff Markdown and human HTML are independently derived from it; neither derived artifact is the other's source.
@@ -66,7 +67,7 @@ profile document.
66
67
  - Schema-v2 final reports author `clarificationItems[]` in data.json; task-specific HTML renders the question and response controls directly from those IDs, and AI handoff Markdown renders the same array as one headed section per row for the next agent. The remaining table-layout rules describe schema-v1 compatibility and analysis-worker result tables only.
67
68
  - **Every row that is still `open` and carries `Blocks=approval` records two more fields.** Withholding approval is the most expensive thing a report does to a run, and until these fields existed a blocker could not be told apart from a question nobody had put to the user.
68
69
  - `origin` — who raised it. `worker-finding` (an analyser or verifier reached it on its own evidence), `material-gap` (neither the brief nor the codebase answers it), or `lead-directed` (the lead's own judgment, **including anything the lead instructed a worker to raise**). A lead that seeds its conclusion into a worker prompt and then reports the worker's agreement as an independent finding has mislabelled the row; that shape is what let one run block on a question its own lead had authored.
69
- - `userConfirmation` — what happened before the row was written. `asked-and-answered`, `asked-awaiting` (asked, no answer yet), or `deferred-no-interactive-session` (this run had no user to ask). An answered question stops being a blocker: record the answer in `userInput`, move `status` to `answered`, and let the plan proceed.
70
+ - `userConfirmation` — what happened before the row was written. `asked-and-answered`, `asked-awaiting` (asked, no answer yet), or `deferred-no-interactive-session` (this run had no user to ask). Record an answer in `userInput` and move `status` to `answered`.
70
71
  - Neither field is required once `status` is `answered` / `resolved` — the record lives in `userInput` by then.
71
72
  - **Legacy canonical column schema (must match `templates/reports/final-report.template.md` §1 exactly):** every `## 1. Clarification Items` table has exactly these 4 columns, in this order:
72
73
  `| <record-meta> | Statement | Expected form | User input |` (the first header is the i18n `columns.recordMeta` label — `Record`).
@@ -10,7 +10,7 @@ are collected and convergence finished. Phase 1-5 do not need it.
10
10
 
11
11
  ## Required deliverable shape (final report, in addition to the standard sections)
12
12
 
13
- - **Plan link & approval evidence**: path to the approved `final-report.md`, the exact quoted approval marker, AND the executed stage number / title quoted from the Stage Map row.
13
+ - **Plan link & approval evidence**: path to the approved `final-report.md`, the exact quoted approval marker, AND the executed stage number / title quoted from the Stage Map row. For a selected-direction plan, also quote `selectedDirectionRef.optionId`, `snapshotPath`, and the validated snapshot digest; for a legacy plan, quote the effective `implementation-option` or the Recommended Option fallback.
14
14
  - **Commit list**: each commit's SHA (or short SHA), message, and the plan step(s) / TDD cycle it satisfies
15
15
  - **Diff summary**: `git diff --stat <base>..HEAD` output, plus a per-file one-line summary of changes
16
16
  - **Out-of-plan edits block**: every file edited that was not in the approved plan's file list, with rationale (empty block is acceptable and preferred)
@@ -42,7 +42,7 @@ are collected and convergence finished. Phase 1-5 do not need it.
42
42
 
43
43
  ## Self-review pass before finalising the report (the Okstra lead runs this; do not delegate it)
44
44
 
45
- 1. **Plan coverage** — every step in the approved plan's recommended option must point to a commit (or an explicit `Skipped: <reason>` entry). List gaps. A `RED:` step and its `GREEN:` step pointing to the same merged commit SHA is NOT a coverage gap — one SHA may be shared by both.
45
+ 1. **Plan coverage** — for a selected-direction plan, every step in the approved `plan-ready` stage must point to a commit (or an explicit `Skipped: <reason>` entry), and the diff must preserve the selected snapshot's mechanism and invariants. For a legacy plan, every step in the effective implementation option (explicit frontmatter value or Recommended Option fallback) must point to a commit or an explicit skip. List gaps. A `RED:` step and its `GREEN:` step pointing to the same merged commit SHA is NOT a coverage gap — one SHA may be shared by both.
46
46
  2. **Evidence completeness** — every `Validation evidence` and `TDD evidence` claim has the actual command line and exit code? No paraphrased "tests pass" without output?
47
47
  3. **Out-of-plan honesty** — files in the diff that are NOT in the plan list must appear in the `Out-of-plan edits` block. Cross-check with `git diff --name-only`.
48
48
  4. **Verifier dissent preserved** — if the verifiers in the resolved roster disagree, the disagreement is visible in the report? Synthesis hides nothing?
@@ -23,10 +23,10 @@
23
23
  - **Falsifiable cause candidates:** every root-cause candidate must include supporting evidence, the strongest falsifying evidence checked, confidence, and the next diagnostic action that would disprove it. A candidate that cannot be falsified is too vague for this phase.
24
24
  - **Graph-aware scope:** a graph edge can explain ordering or duplication, but it is not proof of cause by itself. Cite code/log evidence before claiming an upstream related task caused the current symptom.
25
25
  - **Sharp next diagnostic:** end with the single highest-value diagnostic command, log capture, or file inspection that should happen next, plus the expected signal that would confirm or reject the leading cause.
26
- - **Fix-design boundary:** do not design the implementation fix beyond what is necessary to validate the cause. If the cause is credible, route to `implementation-planning` with the verified evidence; if the cause is still unclear, route to another `error-analysis` run with the next diagnostic.
26
+ - **Fix-design boundary:** do not design the implementation fix beyond what is necessary to validate the cause. If the cause is credible, route to `implementation-option-selection` with the verified evidence; if the cause is still unclear, route to another `error-analysis` run with the next diagnostic.
27
27
  - Structured diagnosis and routing contract:
28
28
  - `errorAnalysis` is the source of truth for reproduction status, `EA-NNN` cause candidates, the sharp next diagnostic, and the next route.
29
- - A route to `implementation-planning` requires a credible leading cause referenced by `routing.leadingCauseId` and `begin-planning` as the direction. A route back to `error-analysis` requires the sharp next diagnostic and `continue-investigation` as the direction.
29
+ - A route to `implementation-option-selection` requires a credible leading cause referenced by `routing.leadingCauseId` and `begin-option-selection` as the direction. A route back to `error-analysis` requires the sharp next diagnostic and `continue-investigation` as the direction.
30
30
  - Structure is enforced by `schemas/final-report-v1.0.schema.json` `$defs.ErrorAnalysis`. Cross-field diagnosis and route semantics are enforced by `validators/validate-run.py::_validate_error_analysis_consistency`.
31
31
  - Primary focus areas:
32
32
  - symptom and trigger clarification
@@ -50,5 +50,5 @@
50
50
  {{INCLUDE:_coverage-critic.md}}
51
51
  - Non-goals:
52
52
  - implementation details unless they are necessary to validate the cause
53
- - **source code edits, builds, migrations, or deployments** — this run produces evidence and cause analysis only; the fix belongs to a later `implementation-planning` run followed by an `implementation` run
53
+ - **source code edits, builds, migrations, or deployments** — this run produces evidence and cause analysis only; the fix belongs to a later `implementation-option-selection`, `implementation-planning`, and `implementation` sequence
54
54
  - this run stays in `error-analysis` regardless of user phrasing — the shared anti-escalation rule applies
@@ -58,12 +58,12 @@
58
58
  - Required deliverable shape (final report, in addition to the standard sections):
59
59
  - **Source Implementation Report(s)** (**Enforced:** `validators/validate-run.py` `_validate_verification_target_match` compares `verificationScope`, `worktreePath`, `implementationBaseRef`, `capturedHeadSha`, and the `stageReports` stage set against the digest-verified `instruction-set/verification-target.md`; a snapshot whose digest no longer checks out is ignored rather than trusted. `verificationScope` in particular gates both stage-group eligibility and release-handoff routing, so it is not the report's to restate): the `VERIFICATION_TARGET` snapshot verbatim — verification scope, worktree path, base/head refs, the list of stages under verification, and one row per stage citing its originating implementation final-report (`report_path` from `consumers.jsonl`; render `(report_path unrecorded)` when absent). Every analyser prompt carries the same compact target identity (`**Verification scope:** / **Worktree:** / **Verification base ref:** / **Verification head ref:** / **Verification target path:** / **Verification target digest:**`) and reads the sidecar on demand for the complete diff stat. A worker that cannot confirm its analysis ran against that worktree's delivered diff MUST record a `tool-failure`.
60
60
  - **Verdict vocabulary**: Section 7 (`Final Verdict`) MUST include a `Verdict Token` field whose value is exactly one of `accepted`, `conditional-accept`, or `blocked`. `conditional-accept` requires an explicit, exhaustive list of conditions; ambiguous verdicts ("looks good", "mostly ready") are not allowed. Each condition MUST be recorded as a row in the **Conditional Acceptance Conditions** deliverable (`id` `CA-NNN`, `condition`, `evidenceRequired`, `blocksReleaseHandoff`). The validator enforces verdict↔deliverable consistency: `accepted` ⇒ zero acceptance blockers, `blocked` ⇒ at least one, `conditional-accept` ⇒ at least one condition, and a `release-handoff` routing recommendation is allowed only when the verdict is `accepted`. **Any Acceptance Blocker therefore forces the verdict off `accepted` (to `conditional-accept` or `blocked`); the gates below cite this rule instead of restating the arithmetic.**
61
- - **Acceptance Blockers block** (under section 4): one row per blocker with `id`, `severity` (`critical` / `major` / `minor`), evidence (file path, log excerpt, or test output), and the recommended follow-up phase (`error-analysis` or `implementation-planning`). Empty block is acceptable and preferred — render the single line `- No acceptance blockers found.`
61
+ - **Acceptance Blockers block** (under section 4): one row per blocker with `id`, `severity` (`critical` / `major` / `minor`), evidence (file path, log excerpt, or test output), and the recommended follow-up phase: `error-analysis` for a cause problem, `implementation-option-selection` for a direction problem, or `implementation-planning` for a detailed-plan problem. Empty block is acceptable and preferred — render the single line `- No acceptance blockers found.`
62
62
  - **Residual Risk block** (under section 4): risks that are not blockers but should be tracked, each with mitigation owner and a trigger that would escalate them to a blocker.
63
63
  - **Validation Evidence**: for every requirement in the originating plan or task brief, cite the artifact (commit SHA, test output, log line, MCP SELECT result) that demonstrates coverage. Paraphrased "verified" claims without an artifact are rejected.
64
64
  - **Read-only command log**: any pre-existing test/validation command touched during this run MUST be listed with its exact command line and one honest status — `executed` (ran; carries its exit code) / `advisory` (external Tier 3 did not PASS; carries observed/expected results and remains user-owned) / `env-unavailable` (should run but cannot in this environment — missing replica DB, container, or service; carries the reason, never a faked pass) / `not-configured` (no such qa-command tier) / `rejected` (a mutating/denied token — skipped, carries the denied token). A check that could not run locally is recorded as `env-unavailable` or `advisory` according to the external QA policy — never silently dropped and never reported as `executed` with an invented exit code. Mutating-command prohibition is the shared read-only boundary (see Non-goals); it is not restated per row.
65
65
  - **Could-not-verify roll-up (§5.8.9)**: the template mechanically aggregates every not-confirmed check into one scannable list — `gap` requirement-coverage rows, `advisory` / `not-configured` / `env-unavailable` / `rejected` command rows, and `blocked` manual tests. You do not hand-author it, but you MUST give those rows their honest status so nothing unverified hides across sections: a check silently recorded as `executed`/`covered` will not surface in the roll-up. This is okstra's answer to "say what could not be verified this run."
66
- - **Routing recommendation**: the next safe phase — one of `release-handoff`, `done`, `error-analysis`, `implementation-planning` — tied to the verdict and blocker list. `release-handoff` is allowed ONLY when the Verdict Token is `accepted`. `release-handoff` is additionally allowed ONLY when the verification scope (the `Verification scope:` line of the injected `VERIFICATION_TARGET` block, recorded as the report's `verificationScope` field) is `whole-task`; a `single-stage` accepted run routes to `release-handoff(stage-group)` (or `implementation` / `done`); plain `release-handoff` remains whole-task-only. Enforcement: `validators/validate-run.py` rejects a `single-stage` report whose routing cites plain `release-handoff`.
66
+ - **Routing recommendation**: the next safe phase — one of `release-handoff`, `done`, `error-analysis`, `implementation-option-selection`, `implementation-planning` — tied to the verdict and blocker list. `release-handoff` is allowed ONLY when the Verdict Token is `accepted`. `release-handoff` is additionally allowed ONLY when the verification scope (the `Verification scope:` line of the injected `VERIFICATION_TARGET` block, recorded as the report's `verificationScope` field) is `whole-task`; a `single-stage` accepted run routes to `release-handoff(stage-group)` (or `implementation` / `done`); plain `release-handoff` remains whole-task-only. Enforcement: `validators/validate-run.py` rejects a `single-stage` report whose routing cites plain `release-handoff`.
67
67
  - **Verified-row recording** (single-stage scope only): when the Verdict Token is `accepted`, the lead MUST run `okstra handoff record-verified --plan-run-root <plan-run-root> --stage <N> --report-path <final-report.md path> --data-json <final-report data.json path>` and quote the command + exit code in the report. The helper re-validates taskType/scope/verdict from data.json, so a non-accepted or whole-task report is rejected at the tool layer. **Enforced:** `validators/validate-run.py` `_validate_verified_row_recorded` requires a `verified` row in `runs/implementation-planning/consumers.jsonl` for every accepted stage — the helper validated its own inputs but nothing checked it had ever run, leaving reports that said `accepted` while the registry said unverified, so the stage was never offered for a stage-group PR.
68
68
  - Clarification request policy (phase-specific addendum — shared policy is in `_common-contract.md`):
69
69
  - populate `## 1. Clarification Items` only when a blocker hinges on information only the user can supply (deployment intent, intended target environment, business-rule interpretation); use `Blocks=next-phase` for items that gate continuing to release-handoff
@@ -77,6 +77,6 @@
77
77
  - **Acceptance critic (opt-in)**: when `convergence.critic.enabled=true` (chosen via the okstra-run picker or `--critic`), a reused-worker **acceptance devil's-advocate** pass is dispatched concurrently with the first convergence reverify round to surface candidate acceptance blockers the verifiers may have missed; candidates are verified only after convergence completes. Each candidate is verified **confirm-or-downgrade**: confirmed → an `Acceptance Blockers` row; unconfirmed → a `Residual Risk` row (never dropped). See `prompts/lead/convergence.md` "Acceptance critic pass (final-verification)".
78
78
  - Non-goals:
79
79
  - proposing unrelated refactors beyond the delivered scope
80
- - **source code edits, follow-up bug fixes, or scope expansion** — this run renders a verdict only; defects detected here become inputs to a new `error-analysis` or `implementation-planning` run
80
+ - **source code edits, follow-up bug fixes, or scope expansion** — this run renders a verdict only; defects detected here become inputs to a new `error-analysis`, `implementation-option-selection`, or `implementation-planning` run according to whether the cause, direction, or detailed plan is invalid
81
81
  - read-only execution of pre-existing test or validation commands is permitted, but any command that mutates source, schema, or deployment state is forbidden
82
82
  - this run records detected issues and ends — the shared anti-escalation rule forbids in-run fixes regardless of user phrasing
@@ -38,6 +38,13 @@
38
38
  "executing builds, migrations, deployments, or any state-mutating command",
39
39
  "starting `implementation-planning` or `implementation` inside this run (each must be a separate run, and `implementation` additionally requires an approved `implementation-planning` deliverable)"
40
40
  ],
41
+ "implementation-option-selection": [
42
+ "source or configuration edits, refactors, or fix attempts",
43
+ "tests, builds, migrations, deployments, or any state-mutating command",
44
+ "detailed file lists, stage maps, execution commands, or plan approval",
45
+ "starting `implementation-planning` or `implementation` inside this run",
46
+ "displaying more than three merged candidates or omitting rejected-candidate audit records"
47
+ ],
41
48
  "implementation-planning": [
42
49
  "source code edits of any kind (Edit/Write on project source files is forbidden)",
43
50
  "file writes outside the run`s artifact directories (`reports/`, `prompts/`, `state/`, `manifests/`, `worker-results/`, `status/`, `sessions/`) and the task-root qa tree (`<task_root>/qa/` — the Tier3 conformance scripts, manifest, and tsconfig this phase MUST write per the Stage Map conformance contract); in particular, do not write to `docs/superpowers/specs/` or `docs/superpowers/plans/`",
@@ -0,0 +1,35 @@
1
+ # Implementation Option Selection Profile
2
+
3
+ - Purpose: compare feasible implementation directions before planning, preserving a read-only record of the evidence and trade-offs that selects the direction to plan
4
+ - Required workers:
5
+ - claude
6
+ - codex
7
+ - antigravity
8
+ - report-writer
9
+ - Optional workers (opt-in via `--workers`):
10
+ - grok
11
+ - kimi
12
+ {{INCLUDE:_common-contract.md}}
13
+ - Brief consumption:
14
+ - Apply the shared reporter-confirmation precondition exactly as written. Unresolved `intent-check:` and `conversion-block:` rows use `Blocks=next-phase`.
15
+ - Treat each stable brief end-state ID as a required evaluation target. A missing ID is a preparation failure; do not invent a replacement requirement.
16
+ - Worker direction-selection procedure:
17
+ - In `candidate-comparison` mode, produce candidate, supporting and contradicting evidence, criterion scores, and requirement mappings.
18
+ - In `candidate-comparison` mode only, submit at most three candidates. A candidate must be feasible from inspected evidence, not from an assumed future change.
19
+ - In `preselected-validation` mode, receive one preselected direction from the lead and validate its evidence, counterevidence, criterion scores, and requirement mappings. The worker must not generate new candidates.
20
+ - Do not produce detailed file lists, stage maps, execution commands, or a plan approval request.
21
+ - Pre-selection context exploration:
22
+ - In `candidate-comparison` mode, inspect the code paths, interfaces, tests, and constraints needed to distinguish candidates before assigning scores.
23
+ - In `preselected-validation` mode, inspect the code paths, interfaces, tests, and constraints needed to validate the one preselected direction.
24
+ - Record uncertainty and contradictory evidence instead of turning it into a candidate preference.
25
+ - Option evaluation rules:
26
+ - `candidate-comparison` generates alternatives. `preselected-validation` validates one preselected direction and does not rank or replace it with an alternative.
27
+ - In `candidate-comparison` mode, the lead merges overlapping candidates, then re-evaluates every merged candidate against the same criteria before ranking it.
28
+ - In `candidate-comparison` mode, display at most three merged candidates and record every rejected candidate with its rejection reason and cited evidence for audit.
29
+ - Map every displayed candidate or preselected direction to the stable brief end-state IDs it satisfies, preserves, or leaves unresolved.
30
+ - Cross-verification mode:
31
+ - Phase 5.5 convergence runs in adversarial mode (`convergence.adversarial=true`).
32
+ - Non-goals:
33
+ - source or configuration edits, tests, builds, migrations, deployments, or other state-mutating commands
34
+ - detailed implementation planning, file-change specifications, stage maps, execution commands, or user approval
35
+ - starting `implementation-planning` or any other lifecycle phase inside this run