okstra 0.179.2 → 0.180.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (78) hide show
  1. package/README.md +1 -1
  2. package/dist/cli-registry.mjs +14 -0
  3. package/dist/cli-registry.mjs.map +1 -1
  4. package/dist/commands/execute/incremental-carry.mjs +9 -8
  5. package/dist/commands/execute/incremental-carry.mjs.map +1 -1
  6. package/dist/commands/execute/plan-verify.mjs +3 -1
  7. package/dist/commands/execute/plan-verify.mjs.map +1 -1
  8. package/dist/commands/report/approval-decision.d.mts +1 -0
  9. package/dist/commands/report/approval-decision.mjs +21 -0
  10. package/dist/commands/report/approval-decision.mjs.map +1 -0
  11. package/dist/commands/report/design-snapshot.d.mts +1 -0
  12. package/dist/commands/report/design-snapshot.mjs +19 -0
  13. package/dist/commands/report/design-snapshot.mjs.map +1 -0
  14. package/docs/architecture/storage-model.md +1 -1
  15. package/docs/architecture.md +10 -10
  16. package/docs/cli.md +11 -8
  17. package/docs/project-structure-overview.md +15 -6
  18. package/docs/task-process/implementation-planning.md +2 -2
  19. package/package.json +1 -1
  20. package/runtime/BUILD.json +2 -2
  21. package/runtime/agents/workers/report-writer-worker.md +15 -164
  22. package/runtime/prompts/launch.template.md +6 -5
  23. package/runtime/prompts/lead/adapters/cmux.md +1 -1
  24. package/runtime/prompts/lead/convergence.md +2 -2
  25. package/runtime/prompts/lead/okstra-lead-contract.md +19 -18
  26. package/runtime/prompts/lead/plan-body-verification.md +39 -18
  27. package/runtime/prompts/lead/report-writer.md +64 -423
  28. package/runtime/prompts/lead/team-contract.md +1 -1
  29. package/runtime/prompts/profiles/_clarification-recommendation.md +5 -4
  30. package/runtime/prompts/profiles/_common-contract.md +3 -3
  31. package/runtime/prompts/profiles/_implementation-deliverable.md +1 -1
  32. package/runtime/prompts/profiles/change-impact-analysis.md +1 -1
  33. package/runtime/prompts/profiles/error-analysis.md +1 -1
  34. package/runtime/prompts/profiles/feature-analysis.md +1 -1
  35. package/runtime/prompts/profiles/implementation-planning.md +13 -11
  36. package/runtime/prompts/profiles/improvement-discovery.md +1 -1
  37. package/runtime/prompts/profiles/project-analysis.md +1 -1
  38. package/runtime/prompts/profiles/requirements-discovery.md +1 -1
  39. package/runtime/python/okstra_ctl/adapters/hosts/claude-code/relay.md +2 -1
  40. package/runtime/python/okstra_ctl/adapters/hosts/external/relay.md +1 -1
  41. package/runtime/python/okstra_ctl/agent_activity.py +23 -3
  42. package/runtime/python/okstra_ctl/agent_prompt_cli.py +6 -6
  43. package/runtime/python/okstra_ctl/analysis_packet.py +43 -2
  44. package/runtime/python/okstra_ctl/approval_decisions.py +327 -0
  45. package/runtime/python/okstra_ctl/design_snapshot.py +134 -0
  46. package/runtime/python/okstra_ctl/dispatch_core.py +62 -4
  47. package/runtime/python/okstra_ctl/dispatch_state.py +29 -4
  48. package/runtime/python/okstra_ctl/execution_mutation_audit.py +6 -2
  49. package/runtime/python/okstra_ctl/final_report_schema.py +24 -15
  50. package/runtime/python/okstra_ctl/incremental_carry.py +128 -16
  51. package/runtime/python/okstra_ctl/incremental_scope.py +4 -1
  52. package/runtime/python/okstra_ctl/path_hints.py +12 -0
  53. package/runtime/python/okstra_ctl/paths.py +12 -0
  54. package/runtime/python/okstra_ctl/plan_items_cli.py +113 -16
  55. package/runtime/python/okstra_ctl/ports/worker_dispatch.py +2 -1
  56. package/runtime/python/okstra_ctl/render.py +48 -1
  57. package/runtime/python/okstra_ctl/render_final_report.py +7 -6
  58. package/runtime/python/okstra_ctl/report_assembly.py +354 -0
  59. package/runtime/python/okstra_ctl/report_contract.py +2 -1
  60. package/runtime/python/okstra_ctl/report_finalize.py +60 -22
  61. package/runtime/python/okstra_ctl/report_inputs.py +72 -0
  62. package/runtime/python/okstra_ctl/report_markdown.py +69 -8
  63. package/runtime/python/okstra_ctl/report_narrative.py +319 -0
  64. package/runtime/python/okstra_ctl/report_projections.py +265 -0
  65. package/runtime/python/okstra_ctl/run.py +25 -9
  66. package/runtime/python/okstra_ctl/schema_excerpt.py +11 -6
  67. package/runtime/python/okstra_ctl/stage_fix_carry.py +4 -4
  68. package/runtime/python/okstra_ctl/stage_ledger.py +132 -18
  69. package/runtime/python/okstra_ctl/stage_map.py +70 -22
  70. package/runtime/python/okstra_ctl/team.py +1 -1
  71. package/runtime/python/okstra_ctl/worker_dispatch.py +5 -2
  72. package/runtime/python/okstra_ctl/worker_prompt_body.py +35 -0
  73. package/runtime/python/okstra_ctl/worker_prompt_policy.py +31 -3
  74. package/runtime/schemas/final-report-v3.0.schema.json +10210 -0
  75. package/runtime/schemas/report-narrative-v3.0.schema.json +30 -0
  76. package/runtime/templates/report-writer-prompt-preamble.md +15 -21
  77. package/runtime/templates/reports/html/macros/forms.html +6 -4
  78. package/runtime/validators/validate-run.py +258 -10
@@ -15,8 +15,8 @@ Plan-body verification runs **after** finding convergence and **after** the repo
15
15
  ```
16
16
  Phase 4 workers produce independent analyses (Findings F-001…)
17
17
  → Phase 5.5 FINDING convergence ([convergence](./convergence.md), sections "Convergence Algorithm" through "Convergence State Artifact")
18
- → Phase 6 report-writer authors final-report data.json (consolidated Option Candidates / Stepwise Execution Order / Dependency / Validation Checklist / Rollback)
19
- → okstra plan-items extract + validate creates the deterministic P-* queue
18
+ → Phase 6 report-writer authors report-writer-narrative Markdown (consolidated Option Candidates / Stepwise Execution Order / Dependency / Validation Checklist / Rollback)
19
+ → okstra plan-items extract --narrative + validate --narrative creates the deterministic P-* queue
20
20
  → PLAN-BODY VERIFICATION ROUND ← this contract
21
21
  → final render/validation
22
22
  → User Approval gate (the frontmatter `approved:` flip is honoured by run-prep only when this round's Gate result is `passed` or `passed-with-dissent`)
@@ -56,9 +56,9 @@ The shared Majority definition and the auto-disable rule (fewer than 2 analyser
56
56
  From the report-writer's draft of `## 5.4 Implementation Plan Deliverables`, the lead creates the verification queue only through this sequence (see also `templates/reports/final-report-v2.template.md` §5.5.9):
57
57
 
58
58
  ```text
59
- okstra plan-items extract --data <data.json> --output <state>/plan-items-....json
59
+ okstra plan-items extract --narrative <report-writer-narrative.md> --output <state>/plan-items-....json
60
60
  → place the persisted `items[]` verbatim in every verifier prompt
61
- → okstra plan-items validate --data <data.json> --items <state>/plan-items-....json
61
+ → okstra plan-items validate --narrative <report-writer-narrative.md> --items <state>/plan-items-....json
62
62
  ```
63
63
 
64
64
  The persisted `items[]` are the sole queue. The lead MUST NOT freely summarise,
@@ -222,9 +222,9 @@ CLI-wrapper calls follow the planned execution surface after
222
222
  consume only `modelExecutionValue`. A missing or invalid invocation contract blocks the
223
223
  round before any host or provider process starts.
224
224
 
225
- 1. Lead runs `okstra plan-items extract --data <data.json> --output <state>/plan-items-....json`, places the persisted `items[]` verbatim in every verifier prompt with the compact `subject` and lossless `payload`, then runs `okstra plan-items validate --data <data.json> --items <state>/plan-items-....json`. Dispatch only after that exact-match validation succeeds.
225
+ 1. Lead runs `okstra plan-items extract --narrative <report-writer-narrative.md> --output <state>/plan-items-....json`, places the persisted `items[]` verbatim in every verifier prompt with the compact `subject` and lossless `payload`, then runs `okstra plan-items validate --narrative <report-writer-narrative.md> --items <state>/plan-items-....json`. Dispatch only after that exact-match validation succeeds.
226
226
 
227
- **Then seed the landing table (BLOCKING):** `okstra plan-items seed --data <data.json>`. `apply-verdicts` in step 8 refuses a verdict whose item has no `planBodyVerification.planItems[]` row, and the report writer leaves that array empty — §5.5.9 is a lead substep that runs after Phase 6 authoring, so nothing before this step has filled it. The seed is idempotent by id and never touches an existing row, so it is safe to re-run between rounds and after a self-fix re-extraction. Skipping it makes step 8 fail with `the report's planBodyVerification has no row for [...]`, which reads as a transcription bug rather than a missing step.
227
+ **Then seed the landing table (BLOCKING):** `okstra plan-items seed --narrative <report-writer-narrative.md> --state <plan-body-verification.json>`. `apply-verdicts` in step 8 refuses a verdict whose item has no `planBodyVerification.planItems[]` row. The report writer never owns that state, so the deterministic seed is the only creator of its rows. The seed is idempotent by id and never touches an existing row, so it is safe to re-run between rounds and after a self-fix re-extraction. Skipping it makes step 8 fail with `plan-body state has no row for [...]`.
228
228
  2. For each analyser worker in the roster (`claude`, `codex`, and `antigravity` if opted in), lead constructs a reverify prompt using the template in §"Plan-body reverify prompt" below.
229
229
  3. Dispatch uses the same wrapper infrastructure as finding convergence, so the `--role-slug` is the same canonical `<role>-worker` that convergence uses — not a round-specific slug. Result file path: `runs/<task-type>/worker-results/<role>-worker-plan-verify-r<N>-implementation-planning-<seq>.md` (e.g. `codex-worker-plan-verify-r1-implementation-planning-003.md`). **`<seq>` is the report's sequence** — the one in this run's `final-report-<task-type>-<seq>` filename, NOT the `workerResults` sequence the initial analysis results carry. The two are equal in most runs and diverge in some (`reports: 004` alongside `workerResults: 005` is a real case), and provenance globs on the report's. Picking the other one makes `_validate_plan_body_verdict_provenance` report that no result file exists while the file is sitting in the directory. The `-worker-` token is load-bearing twice over: §"Plan-body reverify prompt" requires the same anchor headers as convergence, whose `**Audit sidecar path:**` is derived by `okstra_ctl.worker_artifact_paths.audit_sidecar_rel()` inserting `-audit-` after that token — a slug without it makes the header underivable and the helper raises. Record each `planItems[].verdicts[].worker` as the same `<role>-worker` string, because provenance compares it to this filename's prefix. **Enforced:** `tests/contract/test_reverify_dispatch_anchors.py` derives the sidecar from the documented name and re-extracts the prefix the provenance resolver uses.
230
230
  **Verdict provenance (BLOCKING).** Every verdict recorded in `planItems[].verdicts[]` MUST trace back to a dispatch that actually returned a result file at the path above. The whole gate — classification, self-fix eligibility, promotion, `gateBlockedBy` — is computed from these votes, so an unbacked vote lets the round be skipped while the gate still reads `passed`. **Enforced:** `validators/validate-run.py` `_validate_plan_body_verdict_provenance` fails any `verdicts[].worker` with no matching `<worker>-plan-verify-r<N>-<task-type>-<seq>.md` result file. Recording a `verification-error` for a dispatch that produced no result is the correct way to represent a failed worker — inventing an `AGREE` is a contract violation.
@@ -237,7 +237,9 @@ round before any host or provider process starts.
237
237
  - `partial-consensus` — majority `AGREE`, dissenting `DISAGREE` recorded.
238
238
  - `dissent-isolated` — only one worker `DISAGREE`s, others `AGREE` — treat as `partial-consensus` for gate purposes; record dissent. (Distinct from finding-convergence `worker-unique`, which means the *opposite*: only one worker AGREEs. Plan-body classifications use this dedicated label to avoid the collision.)
239
239
  - `majority-disagree` — a *majority* of analysers `DISAGREE` (majority needs ≥2 participating non-error votes; rollback-ordering `DISAGREE(d)` votes are advisory and excluded from the tally), OR any single-vote-blocking kind fires: one `DISAGREE(a)` on any item other than a `P-Var-*` one — where kind `a` never blocks on one vote and takes a majority like `b` / `e` — or one `DISAGREE(f)` on a `P-Req-*` item (see §"Single-vote-blocking kinds"). This classification **blocks approval**.
240
- - `needs-reverify` — a single-vote-blocking kind fired but the item has **fewer than 2 participating non-error votes**, i.e. the lone dissent was never cross-verified because its peer returned `verification-error`. A single-vote-blocking kind means "one *confirmed* DISAGREE is enough"; an unconfirmed one is not, and on a `P-Var-*` item none fires at all — its kind `a` never blocks on one vote and takes a majority like `b` / `e`. This does **not** block approval — blocking on it would make a worker failure produce a stricter gate than a healthy roster, the same paradox the ≥2-vote majority rule already rules out. The item is re-dispatched in the next round (step 7); if it survives the round budget it is promoted per step 8 with a Statement that says verification never completed. **Enforced:** `validators/validate-run.py` `_classify_plan_item_gate` returns `needs-reverify` for this shape and `_recompute_plan_body_gate` folds it into `passed-with-dissent`.
240
+ - `needs-reverify` — one of two shapes the round could not settle.
241
+ - **An even split on a blocking kind.** The majority test is strict, so a panel splitting evenly (1-AGREE / 1-DISAGREE, 2-2, …) reaches neither `full-consensus` nor `majority-disagree`. Until this shape existed it folded into `has-dissent` and the gate passed: two verifiers read the same plan, disagreed on a defect that is not advisory, and the split was recorded and never acted on. An even panel is not only the two-analyser roster — one `UNVERIFIABLE` or one lost dispatch makes any roster even for that item. Re-dispatch those items and record the votes with `--round 2`; a split that survives that round becomes `majority-disagree` and goes to the user, because nothing further is going to settle it. **The round is not optional**: `needs-reverify` folds into `passed-with-dissent`, so without the re-verification this classification would be a label and nothing else. **Enforced:** `validators/validate-run.py` `_validate_unresolved_tie_was_reverified` fails a gate declared over a tie that was never re-verified, and `_classify_plan_item_gate` promotes a tie carrying a round-2 verdict to `majority-disagree`.
242
+ - **A lone dissent nobody cross-verified** — a single-vote-blocking kind fired but the item has **fewer than 2 participating non-error votes**, i.e. the lone dissent was never cross-verified because its peer returned `verification-error`. A single-vote-blocking kind means "one *confirmed* DISAGREE is enough"; an unconfirmed one is not, and on a `P-Var-*` item none fires at all — its kind `a` never blocks on one vote and takes a majority like `b` / `e`. This does **not** block approval — blocking on it would make a worker failure produce a stricter gate than a healthy roster, the same paradox the ≥2-vote majority rule already rules out. The item is re-dispatched in the next round (step 7); if it survives the round budget it is promoted per step 8 with a Statement that says verification never completed. **Enforced:** `validators/validate-run.py` `_classify_plan_item_gate` returns `needs-reverify` for this shape and `_recompute_plan_body_gate` folds it into `passed-with-dissent`.
241
243
  - `contested` only meaningful when `maxRounds > 1`; at default `maxRounds=1`, fold any unresolved item into `partial-consensus`.
242
244
  5. Gate result resolution:
243
245
  - any `majority-disagree` item present AND `gating=true` → `blocked-by-disagreement`
@@ -258,7 +260,7 @@ round before any host or provider process starts.
258
260
  **Record the cause, not just the outcome.** The gate value names the outcome; `planBodyVerification.gateBlockedBy` (array) names every input that blocked it — `majority-disagree`, `coverage-gap`, `non-result`. Two independent inputs can block: a `majority-disagree` plan item, and a Requirement Coverage `gap` / `blocked C-NNN` row (`prompts/profiles/implementation-planning.md` §"Requirement Coverage"). A coverage-only block still renders as `blocked-by-disagreement` because that is the only blocking non-abort value, so **without `gateBlockedBy` the report asserts a worker disagreement that never happened** and the reader hunts for a dissent that does not exist. Leave the array empty for a passing gate. **Enforced:** `validators/validate-run.py` `_validate_gate_blocked_by` cross-checks the declared causes against the recorded verdicts and coverage rows, and fails a passing gate that has a blocking coverage row — the coverage rule was prose-only before.
259
261
 
260
262
  **A coverage row citing this run's own `C-NNN` is not an independent blocker.** When a coverage row's `blocked C-NNN` points at a clarification that step 8 below promoted from a `majority-disagree` item in *this same run*, that blocker is already counted once as the plan item. Counting it again as a coverage gap makes the run block on a clarification it just authored, and the row carries into the next run as a fresh blocker — the Requirement Coverage ↔ Clarification cycle. Such rows are excluded from `coverage-gap`. **Enforced:** `validators/validate-run.py` `_independent_coverage_blockers`.
261
- 6. Lead records `planBodyVerification.participatingAnalysers` as `{rostered, voting}` — how many analysers the roster carried, and how many actually returned a non-error vote. The gate arithmetic is unchanged, but a shrunken roster loosens it silently: with two participating analysers a 1-AGREE / 1-DISAGREE split is a tie, so it never reaches `majority-disagree` and the dissent passes as `dissent-isolated`. A reader comparing two runs' gate values cannot see that without this pair. **Enforced:** `validators/validate-run.py` `_validate_participating_analysers` recomputes `voting` from the recorded verdicts and fails a declared figure the table denies. That pair still counts only *whether* a worker voted: an analyser that answers the same verdict to every item is carried in `voting` as a third opinion while contributing no refutation signal, so the gate reads as a three-way cross-check backed by two. **Enforced (advisory):** `validators/validate-run.py` `_detect_uniform_verifier` reports any worker whose every vote in the round was one verdict, with its item count — it does not fail the run, because a unanimous round is also a legitimate outcome and no ratio separates the two reliably. **Copy every such warning into `planBodyVerification.uniformVerifiers[]`** as `{worker, verdict, itemCount}`; the renderer prints it directly beneath the gate value in both the Markdown and HTML reports. A warning that exists only in the scorer's JSON is not a warning the report's reader ever sees, and the gate line alone reads as a wider cross-check than the round actually was.
263
+ 6. Lead records `planBodyVerification.participatingAnalysers` as `{rostered, voting}` — how many analysers the roster carried, and how many actually returned a non-error vote. The gate arithmetic is unchanged, but a shrunken roster changes what the round can settle: with two participating analysers a 1-AGREE / 1-DISAGREE split is a tie, so it reaches neither consensus nor `majority-disagree` and the item has to go back for a round (see `needs-reverify` above). A reader comparing two runs' gate values cannot see the roster that produced them without this pair. **Enforced:** `validators/validate-run.py` `_validate_participating_analysers` recomputes `voting` from the recorded verdicts and fails a declared figure the table denies. That pair still counts only *whether* a worker voted: an analyser that answers the same verdict to every item is carried in `voting` as a third opinion while contributing no refutation signal, so the gate reads as a three-way cross-check backed by two. **Enforced (advisory):** `validators/validate-run.py` `_detect_uniform_verifier` reports any worker whose every vote in the round was one verdict, with its item count — it does not fail the run, because a unanimous round is also a legitimate outcome and no ratio separates the two reliably. **Copy every such warning into `planBodyVerification.uniformVerifiers[]`** as `{worker, verdict, itemCount}`; the renderer prints it directly beneath the gate value in both the Markdown and HTML reports. A warning that exists only in the scorer's JSON is not a warning the report's reader ever sees, and the gate line alone reads as a wider cross-check than the round actually was.
262
264
 
263
265
  **Check each verifier's verdict distribution before the next round.** Read the `okstra plan-verify` warnings alongside the gate value. Two shapes mean the roster was narrower than it looks: a verifier whose every vote was one token, and a verifier that returned no vote for items it was assigned. Both are contract violations of the adversarial posture, not stylistic preferences — the verifier is told to open the cited evidence and judge it.
264
266
 
@@ -266,7 +268,7 @@ round before any host or provider process starts.
266
268
 
267
269
  **How the corrective round is recorded.** The first prompt was dispatched, so it is immutable — `--replace-undispatched` refuses it, correctly. Materialize the correction under a NEW `--invocation-id` and a new prompt path. Before linking its result, retire the first attempt's link: `okstra agent-prompt reject-result --run-manifest <path> --dispatch-id <first dispatch id> --superseded-by <corrective dispatch id> --reason "<what was wrong with the returned result>"`. Without that step the corrective `link-result` fails with `agent result is already linked to another dispatch`, which is how a worker that ran for twenty minutes and wrote a good result ends up unrecordable. Nothing is deleted: the rejected link stays in `agentResultLinks` carrying `supersededBy` and `rejectionReason`, so the ledger shows both attempts and why the second exists.
268
270
 
269
- Then lead writes `runs/<task-type>/state/plan-body-verification-<task-type>-<seq>.json` (schema below), **appending this round** — one new `roundHistory[]` entry plus this round's votes on each verified item's `planItems[].rounds[]`. The file accumulates across rounds; it is never truncated to the latest one. After `okstra plan-verify` exits 0, lead sets that new round's `completedAt` to the current ISO 8601 UTC time exactly once; a prior round's `completedAt` is immutable. Lead then populates `### 5.5.9 Plan Body Verification` in the final report's data.json (`implementationPlanning.planBodyVerification`, schema `schemas/final-report-v2.0.schema.json`; template at `templates/reports/final-report-v2.template.md`). The §5.5.9 body is **grouped by plan item**: `planItems[]`, each carrying its `id`, its plain-language `subject` (rendered as the item heading), an optional `sourceSection`, an optional `clarificationId` (the `C-<N>` this item blocks on when `majority-disagree`), and a `verdicts[]` list (`worker / verdict / breakageKind / note`) — one verdict row per worker under that item. The renderer prints three fixed legends (gate values, verdict tokens, breakage kinds a–f) so the reader can decode every cell without opening this spec. The older flat `#### Verdict details` table (`Plan item / Worker / …`, one row per plan-item × worker pair) is superseded by the grouped layout — it hid *what* each vote was about behind a bare `P-*` ID; the subject heading is the fix. The validator's `Plan Body Verification` + `Gate result:` substring checks still gate this section.
271
+ Then lead writes `runs/<task-type>/state/plan-body-verification-<task-type>-<seq>.json` (schema below), **appending this round** — one new `roundHistory[]` entry plus this round's votes on each verified item's `planItems[].rounds[]`. The file accumulates across rounds; it is never truncated to the latest one. After `okstra plan-verify` exits 0, lead sets that new round's `completedAt` to the current ISO 8601 UTC time exactly once; a prior round's `completedAt` is immutable. Report assembly later projects the completed nested `planBodyVerification` into the final record. The rendered §5.5.9 body is **grouped by plan item**: `planItems[]`, each carrying its `id`, plain-language `subject`, optional `sourceSection`, optional `clarificationId`, and a `verdicts[]` list — one verdict row per worker under that item. The renderer prints fixed legends for gate values, verdict tokens, and breakage kinds. The older flat `#### Verdict details` table is historical v2 presentation only.
270
272
  7. **Self-fix loop (one rewrite, targeting planner-fixable defects).** After round 1, lead may run one report-writer rewrite when at least one `majority-disagree` item has a majority of its `DISAGREE` verdicts at `fixability == planner-fixable`. The targeted re-verification after that rewrite is round 2. After round 2, stop automatic self-fix regardless of outcome. Classify every remaining item as `user-decision`, `noncritical-dissent`, or `correctness-critical`. A second automatic self-fix is a contract violation. The fixed order is initial verification → one planner self-fix → targeted re-verification → user gate.
271
273
  - **Group the targets by cause before instructing (BLOCKING).** Blocked items are usually several derivatives of one defect — one constant declared twice, one responsibility given two owners — and the coverage rows that cite them fail as a consequence, not independently. Lead MUST partition this round's targets into cause groups and instruct each group as **"remove this cause"**, naming the derivatives it accounts for. **Handing report-writer a bare item list is forbidden**: patched one at a time, each correction leaves the sibling sections still asserting the old value, so the next round re-finds the same family and the budget drains without converging. Record the partition in `planBodyVerification.selfFixGroups[]` (`round`, `causeSummary`, `itemIds`). One group per item is a legitimate outcome only when the items genuinely share no cause — recorded that way, it is a visible diagnosis rather than a skipped one. **Enforced:** `validators/validate-run.py` `_validate_self_fix_grouping` requires the partition, ties `selfFixRoundsApplied` to the highest recorded round, and fails any corrected item that belongs to no group.
272
274
  - lead instructs report-writer to rewrite the items in each cause group (NOT a full draft regeneration; procedure in [report-writer](./report-writer.md) §"Self-fix rewrite").
@@ -274,10 +276,10 @@ round before any host or provider process starts.
274
276
  - **Drop plan items whose element the round deleted.** A self-fix rewrite may remove a plan element (a validation check, a rollback row). `P-*` ids are positional, so a deletion shifts every later row and silently re-points surviving verdicts at their neighbours — and a verdict recorded against a removed element keeps blocking a gate while being unfindable in the plan, so reading the plan never reveals the cause. After each round, re-extract plan items with `okstra plan-items extract` and re-verify any item whose `subject` no longer matches; never carry the old vote forward across a shift. **Enforced:** `validators/validate-run.py` `_validate_verdicts_match_current_subjects` (re-pointing) and `_validate_plan_item_extraction_completeness` (dangling ids).
275
277
  - **Classify each cause group before instructing it (BLOCKING).** A group is either an *authoring* defect — the plan says something wrong, incomplete, or self-contradictory, which self-fix owns — or a *citation* defect, where the plan points at an analysis artifact incorrectly. Only the first is self-fix work. For the second the finding already exists and already went through convergence, so the fix is to re-cite the converged artifact; instructing report-writer to re-derive the fact means the author reads the source material and produces a **finding that never went through convergence**, which the plan then carries as if it had. That is the role boundary the lead contract draws ("keep analysis, execution, verification, and report authoring responsibilities distinct; return defects to the role that owns them"), and report-writer is authoring-only by its own contract. `P-Req-*` items with breakage kind `f` are where this goes wrong most often: the question is usually whether a coverage row points correctly at something already measured, not whether the measurement is right. State the classification in the group's instruction so the author knows which of the two it is being asked to do.
276
278
  - **A verdict older than the last self-fix is not a verdict (BLOCKING).** A verdict cast in round 1 judged the text before the only automatic rewrite. Once that rewrite runs, the judgement is about a plan that no longer exists. `--round <N>` on `apply-verdicts` stamps each row, and `validators/validate-run.py` `_validate_verdict_rounds_outlive_self_fix` fails any non-carried item whose verdict round is at or before `selfFixRoundsApplied`. Before declaring the gate, every item still holding a pre-self-fix verdict MUST be re-verified in round 2.
277
- - lead re-runs plan-body verification (focused on the corrected items + adjacent items the rewrite touched, plus any `needs-reverify` items whose peer failed to vote last round). After re-verification, overwrite `planItems[].verdicts` with the new verdicts. **The round's verdicts MUST be transcribed into `planBodyVerification.planItems[].verdicts` in the final report's data.json before the gate is declared** — the gate is re-derived from that table, so declaring a gate over an empty one leaves it unauditable. **Enforced:** `_validate_round_recorded_verdicts`. Transcribe with `okstra plan-items collect-verdicts --result <worker>=<path> … --items <plan-items.json> --output <verdicts.json>` then `okstra plan-items apply-verdicts --data <data.json> --verdicts <verdicts.json> --round <N>`, never with a per-round script: the CLI reads the response shape this section fixes and **fails** on an assigned item the worker left unanswered, on a verdict for an item outside the queue, and on a `DISAGREE` with no breakage kind. A hand-written regex reports none of those — it drops them, and the round is then scored on a table that silently does not match the queue.
279
+ - lead re-runs plan-body verification (focused on the corrected items + adjacent items the rewrite touched, plus any `needs-reverify` items whose peer failed to vote last round). After re-verification, overwrite `planBodyVerification.planItems[].verdicts` in the convergence-owned state file; the final report does not exist yet. **The round's verdicts MUST be recorded before the gate is declared** — the gate is re-derived from that table, so declaring a gate over an empty one leaves it unauditable. **Enforced:** `_validate_round_recorded_verdicts`. Transcribe with `okstra plan-items collect-verdicts --result <worker>=<path> … --items <plan-items.json> --output <verdicts.json>` then `okstra plan-items apply-verdicts --state <plan-body-verification.json> --verdicts <verdicts.json> --round <N>`. Score the result with `okstra plan-verify --narrative <report-writer-narrative.md> --state <plan-body-verification.json>`. These commands fail on an assigned item the worker left unanswered, on a verdict for an item outside the queue, and on a `DISAGREE` with no breakage kind.
278
280
  - for an item whose `majority-disagree` was resolved by self-fix, record `self-fixed in round <N>: <what was fixed>` in `planItems[].selfFixNote`. A resolved item does not create a clarification.
279
281
  - **Each round is a worker batch.** Before dispatching round N ≥ 2, reclaim the previous round's completed verifiers exactly as at any other batch boundary ([okstra-lead-contract](./okstra-lead-contract.md) "Run-scoped worker-resource lifecycle") and emit `PROGRESS: phase-batch-cleanup panes=<n>`, then announce the round with `PROGRESS: phase-5.5.9-plan-verify round=<N> items=<count>`. Saying a round will "reuse" the previous verifiers and then dispatching under fresh names leaves every prior round holding its panes — five rounds of that is what exhausts the pane budget and blocks the next dispatch. **Enforced:** `validators/validate_session_conformance.py` `_check_plan_verify_cleanup_checkpoints` requires both lines once the state file records two or more rounds.
280
- - **Round completion.** A round is complete only after the renderer has run on the corrected data.json, lead has appended the round to the state file per step 6, lead has reconciled instructed groups against applied corrections — every `itemIds` entry either carries a `selfFixNote` or is still recorded as broken — **`okstra plan-verify --report <report>` exits 0** (step 5), and lead has then set that round's immutable `completedAt`. A round left with a non-zero exit carries its defect into the next round's inputs, which is how a mis-scored gate survives a whole self-fix budget. A round that was instructed but never rendered has not happened, and counting it inflates the budget that gates promotion. The state-file append is not optional bookkeeping: the next re-verification overwrites data.json's `planItems[].verdicts`, so a round that never reached `roundHistory[]` leaves no record anywhere of what it blocked on — which is the whole reason this file exists. **Enforced:** `validators/validate-run.py` `_validate_plan_body_state_rounds` requires one `roundHistory[]` entry per round `1..roundCount`, each carrying its own `gateResult` and cited by at least one item's `rounds[]`, and requires the file's `selfFixRoundsApplied` to match the report's. For a resolved correctness-critical user response, `_validate_target_round_causality` additionally requires the cited round's `completedAt` to be after every linked canonical `user-decision-required` event and no later than every linked canonical `user-decision-evaluated` event.
282
+ - **Round completion.** A round is complete only after the corrected narrative exists, lead has appended the round to the convergence-owned state per step 6, lead has reconciled instructed groups against applied corrections — every `itemIds` entry either carries a `selfFixNote` or is still recorded as broken — **`okstra plan-verify --narrative <report-writer-narrative.md> --state <plan-body-verification.json>` exits 0** (step 5), and lead has then set that round's immutable `completedAt`. A round left with a non-zero exit carries its defect into the next round's inputs. A round that was instructed but never persisted in `roundHistory[]` has not happened. Report assembly and rendering occur only after the round state is terminal. **Enforced:** `validators/validate-run.py` `_validate_plan_body_state_rounds` requires one `roundHistory[]` entry per round `1..roundCount`, each carrying its own `gateResult` and cited by at least one item's `rounds[]`, and requires the state file's `selfFixRoundsApplied` to match the assembled report. For a resolved correctness-critical user response, `_validate_target_round_causality` additionally requires the cited round's `completedAt` to be after every linked canonical `user-decision-required` event and no later than every linked canonical `user-decision-evaluated` event.
281
283
  - **Loop termination.** Lead — not the report-writer worker — records the round count in `planBodyVerification.selfFixRoundsApplied` at each round's end, and why the loop stopped in `planBodyVerification.selfFixStopReason`. The count must equal the highest `round` in `selfFixGroups[]`, so it is derivable from recorded work rather than self-reported. A user-directed correction does not consume the automatic self-fix limit, and a verification failure after that correction does not restart the automatic loop:
282
284
  - `all-resolved` — no planner-fixable `majority-disagree` item remains. Exit.
283
285
  - `no-progress` — the round resolved **zero** planner-fixable items relative to the previous round. Exit even with budget left: the same rewrite would repeat. Newly *introduced* defects count against progress, so a rewrite that trades one defect for another stops the loop rather than churning. A round that re-targets only what the previous round left unresolved is this same conclusion reached one dispatch earlier — exit on it under this reason rather than paying for the round that proves it. **Enforced (advisory):** `validators/validate-run.py` `_detect_self_fix_recurrence` warns on that shape and names this stop reason.
@@ -293,8 +295,8 @@ round before any host or provider process starts.
293
295
  - `Blocks=approval`
294
296
  - the item's `planItems[].clarificationId` set to that `C-<N>` (1:1 link). `validators/validate-run.py` `_validate_plan_body_clarification_matching` recomputes each item's class and fails when a majority-disagree item's `clarificationId` is missing, dangling, or points at a non-`approval` row.
295
297
  - set `approvalContext.classification` to `user-decision` for a majority `needs-user-input` item, `correctness-critical` for `DISAGREE(a)`, `DISAGREE(f)` on `P-Req-*`, or an independent Requirement Coverage blocker, and `noncritical-dissent` for another surviving majority disagreement.
296
- - populate `approvalContext.planItemIds`, `activityIds`, `unblockCondition`, and `recommendedDisposition`. `planItemIds` carries the **extracted item ids verbatim** — the ordinal form the extractor issues (`P-Opt-1`, `P-Step-1.1`), never the human label the plan prose uses for the same thing ("Option C", "Completion B"). The letter label is what that item's `subject` records (§"subject" above: `P-Opt-1` → "Option A: …"), so cite the id and let the subject carry the name; a label written into `planItemIds` reads as an unknown plan item and fails the run. Every option carries a `disposition`: `select` only for `user-decision`, `accept-risk` only for `noncritical-dissent`, and `request-revision` / `reject` for any classification. `correctness-critical` never offers or records `accept-risk`. **Enforced:** `validators/validate-run.py` `_validate_approval_context`.
297
- - **Self-fix exhaustion is not risk acceptance.** A `noncritical-dissent` item remains blocking until the user explicitly selects `accept-risk`. Record the user's non-empty original text and the `user-decision-required` / `user-decision-evaluated` activity references in `approvalContext.resolution`; only then does `validators/validate-run.py` `_resolved_noncritical_dissent_ids` let `_is_dissent_downgraded` fold it into `passed-with-dissent`.
298
+ - record the decision through `okstra approval-decision open`. Each option carries `disposition`, exactly one `reach`, and optional `scopeEffects`. The activity ledger carries affected `planItemIds` and `clarificationRefs`; the approval row never copies those backtrace IDs. `select` is allowed only for `user-decision`, `accept-risk` only for `noncritical-dissent`, and `request-revision` / `reject` for any classification. `correctness-critical` never offers or records `accept-risk`. **Enforced:** `scripts/okstra_ctl/approval_decisions.py` and report assembly.
299
+ - **Self-fix exhaustion is not risk acceptance.** A `noncritical-dissent` item remains blocking until the user explicitly selects `accept-risk`. Record the user's non-empty original text and existing activity check references through `okstra approval-decision resolve`; report assembly derives the report `resolution`.
298
300
  - **Correctness-critical defects cannot be waived.** After the user-directed correction, targeted re-verification of every linked item MUST record only `AGREE` or an acceptable `SUPPLEMENT`, and any independent Requirement Coverage blocker MUST be removed before the row becomes `resolved`. A `DISAGREE` or `verification-error` returns it to `open`. **Enforced:** `validators/validate-run.py` `_validate_correctness_resolution`.
299
301
  - When a correctness-critical `planner-fixable` item is promoted, its `Statement` MUST state "planner self-fix attempted but unresolved" and name the stop reason. `validators/validate-run.py` `_validate_self_fix_before_clarification` fails when a planner-fixable majority item is promoted while the budget is not exhausted — it requires `selfFixRoundsApplied >= 1` **and** `selfFixStopReason` in `{no-progress, max-rounds-reached}`, so neither `all-resolved` nor `not-attempted` can excuse a promotion.
300
302
  - Approval state transitions are fixed:
@@ -303,25 +305,34 @@ round before any host or provider process starts.
303
305
  - `answered → open` when application or checking fails
304
306
  - `open → obsolete` only when a plan change removes the question
305
307
  `open` and `answered` continue to block approval; only `resolved` and `obsolete` are non-blocking. A user-directed correction does not consume the automatic self-fix limit, and a failed check does not restart the automatic loop.
306
- - A terminal row may preserve its original dissent classification only from audited history. Keep superseded votes in `state/plan-body-verification-implementation-planning-<seq>.json`; `validators/validate-run.py` `_historical_plan_item_evidence` recomputes criticality from the recorded `DISAGREE(a|f)` tokens and does not trust the row's classification alone. Every referenced `user-decision-required` / `user-decision-evaluated` activity must cite exactly that row's `C-NNN` and exactly the linked `approvalContext.planItemIds` set. The evaluated activity occurs after every required activity, has `outcome: resolved`, records at least one command whose every `exitCode` is `0`, points `resultPath` at the matching plan-body state artifact, and cites exactly one `plan-body-verification:round-N` evidence token. Round `N` is later than the recorded blocking round; its state votes are all `AGREE` or `SUPPLEMENT`, and they exactly match the final-report verdicts. Its immutable `completedAt` is later than every referenced required activity's canonical event timestamp and no later than every referenced evaluated activity's canonical event timestamp, so an older successful round cannot be relabelled as the response check. This user-response round is recorded in `roundHistory[]` but does not increment `selfFixRoundsApplied` or create another automatic `verification-round-completed` activity. Only the exact round token in `resolution.checkRefs` of a resolved `correctness-critical` row receives that exclusion; the referenced resolved `user-decision-evaluated` activity must match the row's exact `C-NNN` and plan-item set. **Enforced:** `validators/validate-run.py` `_validate_approval_activity_refs`, `_validate_correctness_resolution`, and `_validate_target_round_causality`, plus `validators/validate_session_conformance.py` `_resolved_correctness_reverification_rounds` and `_check_activity_round_counts`. When an independent coverage-only blocker is corrected, keep the `C-NNN` in the now non-blocking Requirement Coverage row's `decisionRefs` and in the matching state-sidecar plan item's `clarificationId`; that item must have no historical blocking dissent and must participate in a round whose `gateBlockedBy` contains `coverage-gap`. A run-wide `coverage-gap` without this item-level `C-NNN` link cannot classify another row. `obsolete` is valid only after current evidence shows that the question or blocker disappeared, or the linked item is historical and removed; a current linked item remains active even when audited history preserves an older classification.
308
+ - A terminal row preserves its original dissent classification only from the convergence-owned state history. Every `user-decision-required` / `user-decision-evaluated` activity cites the row's `C-NNN` in `clarificationRefs` and affected plan items in `planItemIds`. A resolved decision names only existing `A-NNN` checks. Report assembly validates those links and derives the report backtraces; it does not accept copied IDs from the approval ledger. When an independent coverage-only blocker is corrected, keep the `C-NNN` in the non-blocking Requirement Coverage row's `decisionRefs`. `obsolete` is valid only after current evidence shows that the question or blocker disappeared.
307
309
  9. Approval lives in the report record `frontmatter.approved` field — there is no in-body marker line. The user may set it to `true` (via `--approve` or the in-session wizard) only when the Gate result is `passed` or `passed-with-dissent`. **Enforced:** run-prep (`scripts/okstra_ctl/run.py` `_validate_approved_plan`) fail-closes an `approved: true` plan whose record carries a blocking `gateResult` or an open/answered `Blocks: approval` clarification row, and `validators/validate-run.py` `_validate_plan_body_gate_recompute` rejects a declared `gateResult` healthier than the recorded votes.
308
310
 
309
311
  ## `plan-body-verification-<task-type>-<seq>.json` schema
310
312
 
311
- **Which file is authoritative for what.** These two representations are *not* required to agree, and a lead that tries to make them match is doing unnecessary work:
313
+ **Which file is authoritative for what.** Contract v3 keeps both views in one convergence-owned state file before publication:
312
314
 
313
315
  | | records | what a self-fix round does to it |
314
316
  |---|---|---|
315
- | `state/plan-body-verification-<task-type>-<seq>.json` | the round-by-round history, including rounds later superseded | **appends** — a new `roundHistory[]` entry and new `planItems[].rounds[]` votes; earlier rounds are left untouched, and this is the only place they survive |
316
- | `data.json` `implementationPlanning.planBodyVerification` | the **final** state after the self-fix loop | **overwrites** — each re-verification replaces `planItems[].verdicts` |
317
+ | `state/plan-body-verification-<task-type>-<seq>.json` top-level `planItems[]` / `roundHistory[]` | round-by-round history, including superseded rounds | **appends** — earlier rounds remain immutable |
318
+ | the same file's `planBodyVerification` | final state after the self-fix loop | **overwrites** — each re-verification replaces current `planItems[].verdicts` |
319
+ | `data.json` `implementationPlanning.planBodyVerification` | published projection | report assembly copies the validated final state once |
317
320
 
318
- **The gate is computed from data.json.** The state file is the audit trail: after a successful self-fix loop the two legitimately differ (the sidecar shows what *each* round blocked on, data.json shows only the resolved result), and that difference is the record of the fix working. **Enforced:** `validators/validate-run.py` `_validate_plan_body_state_file` requires the file to exist with `schemaVersion` / `planItems` / `roundHistory` once a round has run, then `_validate_plan_body_state_rounds` checks the round coverage described in §"Round protocol" step 7. Neither compares a gate value to data.json — the two views are supposed to differ, and demanding equality would fail every run whose self-fix loop worked.
321
+ **The gate is computed from the nested final projection.** The top-level audit history may differ because it preserves superseded rounds. Report assembly copies the final projection rather than asking the report writer to transcribe it. **Enforced:** `validators/validate-run.py` `_validate_plan_body_state_file` requires the audit keys once a round has run, and `scripts/okstra_ctl/report_assembly.py` reads only the convergence-owned `planBodyVerification` projection.
319
322
 
320
323
  The per-round structures mirror the finding-convergence state artifact ([convergence](./convergence.md) §"Convergence State Artifact"): `roundHistory[]` is the round-level ledger, and each item's `rounds[]` is its per-round vote history — the same split as that file's `roundHistory[]` / `findings[].rounds[]`.
321
324
 
322
325
  ```json
323
326
  {
324
327
  "schemaVersion": "1.1",
328
+ "owner": "convergence",
329
+ "planBodyVerification": {
330
+ "roundCount": 2,
331
+ "gateResult": "passed-with-dissent",
332
+ "gateBlockedBy": [],
333
+ "planItems": [],
334
+ "dissentLog": []
335
+ },
325
336
  "phase": "implementation-planning",
326
337
  "effectiveMaxRounds": 1,
327
338
  "gating": true,
@@ -509,6 +520,11 @@ rerun a free-form requirements analysis. Every DISAGREE includes Fixability.
509
520
 
510
521
  ## Response format
511
522
 
523
+ Head each block with `### <plan item id>` at exactly three hashes. The heading
524
+ depth is not style: the collector parses `^### ` and nothing else, so a block
525
+ written at any other depth is not an unparsed block — it is a verdict that was
526
+ never recorded, and the round is scored on the items that remain.
527
+
512
528
  ### P-Step-3
513
529
  **Verdict**: AGREE | DISAGREE(<a|b|c|d|e|f>) | SUPPLEMENT | UNVERIFIABLE
514
530
  **Fixability** (only when DISAGREE): planner-fixable | needs-user-input — "planner-fixable if it can be fixed using only the code + this plan + the brief". Fixability is not applicable for `UNVERIFIABLE`.
@@ -525,6 +541,11 @@ rerun a free-form requirements analysis. Every DISAGREE includes Fixability.
525
541
  **Explanation**: <2-3 sentences applying the disposition-specific rule>
526
542
  ````
527
543
 
544
+ **This template is the round-1 prompt.** Before rendering a round 2 or later,
545
+ read §"Re-verification rounds (round 2+)" below first — such a round carries
546
+ blocks this template does not have, and re-rendering this one alone is a
547
+ contract violation.
548
+
528
549
  When `config.adversarial == true`, the lead prepends the adversarial framing from §"Adversarial plan-body posture" to the `## Instructions` block: the burden of proof is on the plan, the verifier opens and confirms every accessible cited path / command, and evidence that was opened but is insufficient yields the applicable `DISAGREE(<kind>)` rather than `AGREE`. Inability to inspect because of capability, credential, network, or service state yields `UNVERIFIABLE`, not DISAGREE. The verdict tokens, breakage kinds (a–f), classification, and the majority gate threshold are unchanged. This prepended framing supersedes the template's "Judge solely from plan internal consistency" instruction for the adversarial round.
529
550
 
530
551
  The "Reverify prompt: required-reading suppression" rule in [convergence](./convergence.md) (lightweight mode does NOT inject a `[Required reading]` clause) applies here as well.