okstra 0.180.0 → 0.184.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +1 -1
- package/dist/cli-registry.mjs +16 -2
- package/dist/cli-registry.mjs.map +1 -1
- package/dist/commands/execute/render-bundle.d.mts +4 -2
- package/dist/commands/execute/render-bundle.mjs +46 -5
- package/dist/commands/execute/render-bundle.mjs.map +1 -1
- package/dist/commands/execute/run.mjs +11 -3
- package/dist/commands/execute/run.mjs.map +1 -1
- package/dist/commands/inspect/model-io.d.mts +1 -0
- package/dist/commands/inspect/model-io.mjs +25 -0
- package/dist/commands/inspect/model-io.mjs.map +1 -0
- package/dist/commands/inspect/stage-map.mjs +29 -8
- package/dist/commands/inspect/stage-map.mjs.map +1 -1
- package/dist/commands/inspect/task-list.mjs +52 -6
- package/dist/commands/inspect/task-list.mjs.map +1 -1
- package/dist/commands/inspect/user-response.mjs +14 -4
- package/dist/commands/inspect/user-response.mjs.map +1 -1
- package/dist/commands/lifecycle/check-project.d.mts +1 -0
- package/dist/commands/lifecycle/check-project.mjs +69 -50
- package/dist/commands/lifecycle/check-project.mjs.map +1 -1
- package/dist/commands/lifecycle/contract-check.d.mts +1 -0
- package/dist/commands/lifecycle/contract-check.mjs +18 -0
- package/dist/commands/lifecycle/contract-check.mjs.map +1 -0
- package/dist/commands/lifecycle/preflight.mjs +154 -51
- package/dist/commands/lifecycle/preflight.mjs.map +1 -1
- package/dist/commands/pr/pr.d.mts +1 -0
- package/dist/commands/pr/pr.mjs +19 -1
- package/dist/commands/pr/pr.mjs.map +1 -1
- package/dist/commands/report/agent-activity.mjs +2 -2
- package/dist/commands/report/translate.mjs +3 -0
- package/dist/commands/report/translate.mjs.map +1 -1
- package/dist/lib/host-registry-client.mjs +13 -9
- package/dist/lib/host-registry-client.mjs.map +1 -1
- package/docs/architecture.md +13 -2
- package/docs/cli.md +31 -15
- package/docs/container.md +6 -4
- package/docs/contributor-change-matrix.md +1 -1
- package/docs/for-ai/README.md +2 -2
- package/docs/for-ai/skills/okstra-brief-gen.md +5 -3
- package/docs/for-ai/skills/okstra-code-review.md +4 -4
- package/docs/for-ai/skills/okstra-container-build.md +20 -17
- package/docs/for-ai/skills/okstra-inspect.md +20 -23
- package/docs/for-ai/skills/okstra-manager.md +19 -18
- package/docs/for-ai/skills/okstra-memory.md +2 -2
- package/docs/for-ai/skills/okstra-pr-gen.md +3 -3
- package/docs/for-ai/skills/okstra-rollup.md +14 -13
- package/docs/for-ai/skills/okstra-run.md +7 -3
- package/docs/for-ai/skills/okstra-schedule-gen.md +15 -18
- package/docs/for-ai/skills/okstra-setup.md +7 -7
- package/docs/for-ai/skills/okstra-usage.md +5 -4
- package/docs/for-ai/skills/okstra-user-response.md +50 -32
- package/docs/project-structure-overview.md +30 -27
- package/docs/task-process/README.md +1 -1
- package/docs/task-process/common-flow.md +2 -3
- package/docs/task-process/error-analysis.md +3 -4
- package/docs/task-process/final-verification.md +2 -3
- package/docs/task-process/implementation-planning.md +2 -3
- package/docs/task-process/implementation.md +9 -7
- package/docs/task-process/release-handoff.md +3 -4
- package/docs/task-process/requirements-discovery.md +3 -4
- package/package.json +1 -1
- package/runtime/BUILD.json +2 -2
- package/runtime/agents/workers/claude-worker.md +4 -4
- package/runtime/agents/workers/report-writer-worker.md +3 -3
- package/runtime/agents/workers/translator-worker.md +5 -13
- package/runtime/bin/okstra-error-log.py +51 -11
- package/runtime/bin/okstra-report-translate.py +210 -23
- package/runtime/prompts/host-orchestration/implementation.md +1 -1
- package/runtime/prompts/launch.template.md +11 -14
- package/runtime/prompts/lead/context-loader.md +41 -141
- package/runtime/prompts/lead/convergence.md +8 -6
- package/runtime/prompts/lead/okstra-lead-contract.md +26 -36
- package/runtime/prompts/lead/plan-body-verification.md +211 -30
- package/runtime/prompts/lead/report-writer.md +23 -4
- package/runtime/prompts/lead/team-contract.md +8 -53
- package/runtime/prompts/profiles/_coding-conventions-preflight.md +3 -2
- package/runtime/prompts/profiles/_common-contract.md +1 -1
- package/runtime/prompts/profiles/_implementation-diff-review.md +1 -1
- package/runtime/prompts/profiles/_implementation-executor.md +1 -0
- package/runtime/prompts/profiles/_implementation-verifier.md +4 -4
- package/runtime/prompts/profiles/final-verification.md +1 -1
- package/runtime/prompts/profiles/implementation-planning.md +17 -13
- package/runtime/prompts/profiles/release-handoff.md +0 -1
- package/runtime/prompts/wizard/prompts.ko.json +7 -11
- package/runtime/python/okstra_ctl/adapters/hosts/capability_adapter.py +69 -17
- package/runtime/python/okstra_ctl/adapters/hosts/claude-code/adapter.py +13 -4
- package/runtime/python/okstra_ctl/adapters/hosts/claude-code/relay.md +6 -1
- package/runtime/python/okstra_ctl/adapters/hosts/codex/adapter.py +2 -2
- package/runtime/python/okstra_ctl/adapters/hosts/codex/relay.md +50 -5
- package/runtime/python/okstra_ctl/adapters/hosts/grok/adapter.py +2 -2
- package/runtime/python/okstra_ctl/adapters/hosts/grok/relay.md +66 -5
- package/runtime/python/okstra_ctl/adapters/providers/grok/adapter.py +70 -2
- package/runtime/python/okstra_ctl/agent_activity.py +118 -35
- package/runtime/python/okstra_ctl/agent_invocation.py +19 -6
- package/runtime/python/okstra_ctl/agent_prompt_cli.py +65 -18
- package/runtime/python/okstra_ctl/analysis_inputs.py +5 -4
- package/runtime/python/okstra_ctl/analysis_packet.py +81 -1
- package/runtime/python/okstra_ctl/approval_decisions.py +3 -2
- package/runtime/python/okstra_ctl/attempt_evidence.py +2 -2
- package/runtime/python/okstra_ctl/backfill.py +13 -10
- package/runtime/python/okstra_ctl/batch.py +2 -4
- package/runtime/python/okstra_ctl/build_tools.py +6 -3
- package/runtime/python/okstra_ctl/claim_reproduction.py +101 -0
- package/runtime/python/okstra_ctl/clarification_items.py +27 -13
- package/runtime/python/okstra_ctl/cmux.py +130 -52
- package/runtime/python/okstra_ctl/code_review_target.py +34 -8
- package/runtime/python/okstra_ctl/conformance.py +37 -1
- package/runtime/python/okstra_ctl/consumers.py +5 -4
- package/runtime/python/okstra_ctl/container.py +103 -8
- package/runtime/python/okstra_ctl/context_cost.py +2 -1
- package/runtime/python/okstra_ctl/contract_graph.py +497 -0
- package/runtime/python/okstra_ctl/contract_graph_cli.py +62 -0
- package/runtime/python/okstra_ctl/convergence.py +338 -17
- package/runtime/python/okstra_ctl/convergence_engine.py +10 -18
- package/runtime/python/okstra_ctl/convergence_provenance.py +58 -8
- package/runtime/python/okstra_ctl/convergence_store.py +55 -34
- package/runtime/python/okstra_ctl/design_prep.py +7 -4
- package/runtime/python/okstra_ctl/dispatch_core.py +35 -65
- package/runtime/python/okstra_ctl/dispatch_state.py +134 -59
- package/runtime/python/okstra_ctl/doctor.py +6 -3
- package/runtime/python/okstra_ctl/domain/worker_presentation.py +70 -9
- package/runtime/python/okstra_ctl/entrypoints/hosts.py +16 -30
- package/runtime/python/okstra_ctl/error_log_write.py +35 -30
- package/runtime/python/okstra_ctl/error_report.py +26 -1
- package/runtime/python/okstra_ctl/error_zip.py +27 -5
- package/runtime/python/okstra_ctl/execution_identity.py +3 -2
- package/runtime/python/okstra_ctl/execution_manifest.py +7 -4
- package/runtime/python/okstra_ctl/final_report_schema.py +2 -2
- package/runtime/python/okstra_ctl/fix_cycles.py +2 -2
- package/runtime/python/okstra_ctl/fixed_text.py +39 -0
- package/runtime/python/okstra_ctl/git_reconcile.py +41 -9
- package/runtime/python/okstra_ctl/handoff.py +5 -4
- package/runtime/python/okstra_ctl/i18n.py +4 -2
- package/runtime/python/okstra_ctl/implementation_direction.py +22 -14
- package/runtime/python/okstra_ctl/implementation_outcome.py +4 -7
- package/runtime/python/okstra_ctl/incremental_carry.py +2 -1
- package/runtime/python/okstra_ctl/incremental_scope.py +89 -39
- package/runtime/python/okstra_ctl/index.py +8 -11
- package/runtime/python/okstra_ctl/initial_prompt_materialization.py +79 -7
- package/runtime/python/okstra_ctl/invocation.py +3 -6
- package/runtime/python/okstra_ctl/json_boundary.py +366 -0
- package/runtime/python/okstra_ctl/json_registry.py +10 -12
- package/runtime/python/okstra_ctl/jsonl.py +19 -2
- package/runtime/python/okstra_ctl/lead_events.py +33 -1
- package/runtime/python/okstra_ctl/listing.py +3 -3
- package/runtime/python/okstra_ctl/log_report.py +24 -2
- package/runtime/python/okstra_ctl/manager_cli.py +92 -7
- package/runtime/python/okstra_ctl/manager_store.py +12 -10
- package/runtime/python/okstra_ctl/material.py +5 -1
- package/runtime/python/okstra_ctl/migrate.py +29 -25
- package/runtime/python/okstra_ctl/model_cli.py +3 -15
- package/runtime/python/okstra_ctl/model_io_cli.py +1051 -0
- package/runtime/python/okstra_ctl/mutation_probe.py +13 -4
- package/runtime/python/okstra_ctl/pane_reclaim.py +3 -2
- package/runtime/python/okstra_ctl/paths.py +9 -0
- package/runtime/python/okstra_ctl/plan_items.py +525 -5
- package/runtime/python/okstra_ctl/plan_items_cli.py +842 -32
- package/runtime/python/okstra_ctl/pr_template.py +3 -2
- package/runtime/python/okstra_ctl/project_meta.py +5 -7
- package/runtime/python/okstra_ctl/recap.py +5 -4
- package/runtime/python/okstra_ctl/reconcile.py +21 -27
- package/runtime/python/okstra_ctl/registry/host_discovery.py +3 -2
- package/runtime/python/okstra_ctl/registry/provider_registry.py +3 -2
- package/runtime/python/okstra_ctl/render.py +30 -15
- package/runtime/python/okstra_ctl/render_final_report.py +3 -2
- package/runtime/python/okstra_ctl/report_assembly.py +172 -17
- package/runtime/python/okstra_ctl/report_finalize.py +7 -10
- package/runtime/python/okstra_ctl/report_html/render.py +3 -2
- package/runtime/python/okstra_ctl/report_language.py +3 -2
- package/runtime/python/okstra_ctl/report_markdown.py +13 -1
- package/runtime/python/okstra_ctl/report_narrative.py +40 -8
- package/runtime/python/okstra_ctl/report_synthesis_packet.py +518 -0
- package/runtime/python/okstra_ctl/report_views.py +3 -2
- package/runtime/python/okstra_ctl/rollup.py +65 -4
- package/runtime/python/okstra_ctl/run.py +159 -56
- package/runtime/python/okstra_ctl/run_audit.py +3 -2
- package/runtime/python/okstra_ctl/run_context.py +6 -9
- package/runtime/python/okstra_ctl/run_index_row.py +2 -8
- package/runtime/python/okstra_ctl/schedule_semantics.py +5 -2
- package/runtime/python/okstra_ctl/schema_excerpt.py +4 -2
- package/runtime/python/okstra_ctl/session_transcript.py +27 -1
- package/runtime/python/okstra_ctl/set_work_status.py +64 -38
- package/runtime/python/okstra_ctl/stage_fix_carry.py +4 -2
- package/runtime/python/okstra_ctl/stage_map.py +26 -6
- package/runtime/python/okstra_ctl/stage_targets.py +3 -4
- package/runtime/python/okstra_ctl/team.py +2 -1
- package/runtime/python/okstra_ctl/team_reconcile.py +11 -2
- package/runtime/python/okstra_ctl/time_report.py +51 -4
- package/runtime/python/okstra_ctl/usage_identity.py +2 -1
- package/runtime/python/okstra_ctl/usage_report.py +58 -4
- package/runtime/python/okstra_ctl/user_response.py +1431 -66
- package/runtime/python/okstra_ctl/wizard.py +50 -117
- package/runtime/python/okstra_ctl/work_categories.py +3 -2
- package/runtime/python/okstra_ctl/worker_prompt_body.py +18 -7
- package/runtime/python/okstra_ctl/worker_prompt_contract.py +3 -2
- package/runtime/python/okstra_ctl/worker_runner.py +14 -12
- package/runtime/python/okstra_ctl/workflow.py +2 -1
- package/runtime/python/okstra_ctl/worktree.py +3 -2
- package/runtime/python/okstra_ctl/wrapper_status.py +4 -2
- package/runtime/python/okstra_ctl/write_policy.py +4 -2
- package/runtime/python/okstra_token_usage/antigravity.py +39 -12
- package/runtime/python/okstra_token_usage/collect.py +90 -38
- package/runtime/python/okstra_token_usage/grok.py +127 -0
- package/runtime/schemas/final-report-v2.0.schema.json +21 -0
- package/runtime/schemas/final-report-v3.0.schema.json +21 -0
- package/runtime/schemas/report-synthesis-packet-v1.0.schema.json +140 -0
- package/runtime/skills/okstra-brief-gen/SKILL.md +9 -7
- package/runtime/skills/okstra-code-review/SKILL.md +21 -11
- package/runtime/skills/okstra-container-build/SKILL.md +18 -18
- package/runtime/skills/okstra-inspect/SKILL.md +12 -11
- package/runtime/skills/okstra-inspect/facets/error-zip.md +8 -8
- package/runtime/skills/okstra-inspect/facets/errors.md +2 -2
- package/runtime/skills/okstra-inspect/facets/history.md +9 -14
- package/runtime/skills/okstra-inspect/facets/logs.md +2 -2
- package/runtime/skills/okstra-inspect/facets/recap.md +5 -5
- package/runtime/skills/okstra-inspect/facets/report.md +6 -10
- package/runtime/skills/okstra-inspect/facets/status.md +9 -8
- package/runtime/skills/okstra-inspect/facets/time.md +3 -3
- package/runtime/skills/okstra-manager/SKILL.md +16 -14
- package/runtime/skills/okstra-memory/SKILL.md +3 -3
- package/runtime/skills/okstra-pr-gen/SKILL.md +5 -4
- package/runtime/skills/okstra-rollup/SKILL.md +6 -16
- package/runtime/skills/okstra-run/SKILL.md +9 -9
- package/runtime/skills/okstra-schedule-gen/SKILL.md +21 -17
- package/runtime/skills/okstra-setup/SKILL.md +21 -13
- package/runtime/skills/okstra-setup/references/project-config.md +2 -2
- package/runtime/skills/okstra-usage/SKILL.md +10 -10
- package/runtime/skills/okstra-user-response/SKILL.md +78 -107
- package/runtime/templates/report-writer-prompt-preamble.md +17 -1
- package/runtime/templates/reports/schedule.template.md +4 -4
- package/runtime/templates/worker-error-contract.md +17 -29
- package/runtime/validators/validate-run.py +527 -113
- package/runtime/validators/validate_session_conformance.py +65 -10
|
@@ -16,7 +16,7 @@ Plan-body verification runs **after** finding convergence and **after** the repo
|
|
|
16
16
|
Phase 4 workers produce independent analyses (Findings F-001…)
|
|
17
17
|
→ Phase 5.5 FINDING convergence ([convergence](./convergence.md), sections "Convergence Algorithm" through "Convergence State Artifact")
|
|
18
18
|
→ Phase 6 report-writer authors report-writer-narrative Markdown (consolidated Option Candidates / Stepwise Execution Order / Dependency / Validation Checklist / Rollback)
|
|
19
|
-
→ okstra plan-items
|
|
19
|
+
→ okstra plan-items prepare + prompt + validate-prepared creates and projects the deterministic P-* queue
|
|
20
20
|
→ PLAN-BODY VERIFICATION ROUND ← this contract
|
|
21
21
|
→ final render/validation
|
|
22
22
|
→ User Approval gate (the frontmatter `approved:` flip is honoured by run-prep only when this round's Gate result is `passed` or `passed-with-dissent`)
|
|
@@ -45,9 +45,9 @@ Plan-body verification is configured under `convergence.planBodyVerification` in
|
|
|
45
45
|
| `enabled` | `true` | If `false`, the round is skipped and the approval gate is not blocked by this round (legacy behaviour). |
|
|
46
46
|
| `maxRounds` | `1` | Upper bound. Plan-body verification is consistency / completeness checking, not fact checking — additional rounds rarely help. Range 1–3. |
|
|
47
47
|
| `selfFixMaxRounds` | `1` | One report-writer rewrite at most. The initial verification is round 1; targeted re-verification is round 2 after that rewrite. |
|
|
48
|
-
| `gating` | `true` | If `true` (default), `majority-disagree` blocks approval. If `false`, the round is advisory-only and never blocks approval. |
|
|
48
|
+
| `gating` | `true` | If `true` (default), `majority-disagree` blocks approval. If `false`, the round is advisory-only and never blocks approval. Prepare emits `true` because the plan does not exist yet. After the report-writer draft, `okstra plan-items prepare` (and `seed`) flip it to `false` when `designPreparation.mode` is `no-design-inputs` and the Stage Map has exactly one row. That path keeps extraction and one verification round and does not run the self-fix loop or a sweep batch. Two-or-more stages, a PREP item, or non-empty `designPreparation.items` keep `gating=true`. `--no-plan-verification` is the separate manual opt-out (`enabled=false`). **Enforced:** `okstra_ctl.plan_items.advisory_plan_body_gating`, `validators/validate-run.py` `_validate_advisory_plan_body_gating`. |
|
|
49
49
|
|
|
50
|
-
Default values are emitted into the manifest by `scripts/okstra_ctl/render.py` (`_build_convergence_block`). The ctx knob `OKSTRA_PLAN_VERIFICATION=false` flips `planBodyVerification.enabled` to false.
|
|
50
|
+
Default values are emitted into the manifest by `scripts/okstra_ctl/render.py` (`_build_convergence_block`). The ctx knob `OKSTRA_PLAN_VERIFICATION=false` flips `planBodyVerification.enabled` to false. `gating=false` is not that opt-out: extraction and one round still run.
|
|
51
51
|
|
|
52
52
|
The shared Majority definition and the auto-disable rule (fewer than 2 analyser workers → advisory `gating=false` path) are owned by [convergence](./convergence.md) §"Convergence Algorithm" / §"Configuration" and apply here unchanged.
|
|
53
53
|
|
|
@@ -56,9 +56,9 @@ The shared Majority definition and the auto-disable rule (fewer than 2 analyser
|
|
|
56
56
|
From the report-writer's draft of `## 5.4 Implementation Plan Deliverables`, the lead creates the verification queue only through this sequence (see also `templates/reports/final-report-v2.template.md` §5.5.9):
|
|
57
57
|
|
|
58
58
|
```text
|
|
59
|
-
okstra plan-items
|
|
60
|
-
→ place
|
|
61
|
-
→ okstra plan-items validate --narrative <report-writer-narrative.md> --
|
|
59
|
+
okstra plan-items prepare --narrative <report-writer-narrative.md> --run-manifest <run-manifest>
|
|
60
|
+
→ place `okstra plan-items prompt --run-manifest <run-manifest>` output verbatim in every verifier prompt
|
|
61
|
+
→ okstra plan-items validate-prepared --narrative <report-writer-narrative.md> --run-manifest <run-manifest>
|
|
62
62
|
```
|
|
63
63
|
|
|
64
64
|
The persisted `items[]` are the sole queue. The lead MUST NOT freely summarise,
|
|
@@ -103,6 +103,130 @@ For every detector-produced `(stage, kind)`, extract exactly one plan item named
|
|
|
103
103
|
|
|
104
104
|
When extracting each item, lead also captures a **`subject`** — a plain one-line label (≤12 words) describing *what that item is* in the reader's terms, e.g. `P-Opt-1` → "Option A: split upload v2 into a new module", `P-Step-1.1` → "Stage 1 Step 2: regression-check with `npm run test:v2`". This is a label-capture, not new analysis. The `subject` is what §5.5.9 renders as the per-item heading so the reader knows *what* each AGREE/DISAGREE is about without cross-referencing §4.5; a bare `P-*` ID with no subject is a contract violation. **Enforced:** `validators/validate-run.py` `_validate_plan_item_subject_substance` fails a subject that is a placeholder — under 3 chars, equal to the item id, or shaped like a bare `P-*` id.
|
|
105
105
|
|
|
106
|
+
## Fact and judgement (BLOCKING)
|
|
107
|
+
|
|
108
|
+
A single-vote block has a reason behind it: two spelled-out references
|
|
109
|
+
contradicting each other is a fact one verifier can settle by looking, and a
|
|
110
|
+
fact must not be outvoted. What was missing is that the fact was never checked —
|
|
111
|
+
writing "this path does not exist" was enough to block.
|
|
112
|
+
|
|
113
|
+
So a DISAGREE on a single-vote kind declares what sort of claim it is:
|
|
114
|
+
|
|
115
|
+
- **`claimKind: fact`** — okstra can reproduce it. Carry a `reproduction` probe:
|
|
116
|
+
`path-exists` / `path-absent`, `literal-present` / `literal-absent`, or
|
|
117
|
+
`citations-differ`. **okstra runs it and writes `reproductionResult`; do not
|
|
118
|
+
write that field yourself.** A verifier that reports its own result puts the
|
|
119
|
+
gate on a self-report, and the same claim then settles differently depending
|
|
120
|
+
on who raised it.
|
|
121
|
+
- **`claimKind: judgement`** — no mechanical check exists. It takes a quorum,
|
|
122
|
+
exactly like `b` / `c` / `e`.
|
|
123
|
+
|
|
124
|
+
A claim that cannot be written as one of those three probes is a `judgement`.
|
|
125
|
+
That is a classification, not a demotion: giving a single-vote block to something
|
|
126
|
+
no machine can confirm is what was wrong in the first place.
|
|
127
|
+
|
|
128
|
+
| Declared | Reproduced | Effect |
|
|
129
|
+
|---|---|---|
|
|
130
|
+
| `fact` | `reproduced` | blocks on one vote |
|
|
131
|
+
| `fact` | `not-reproduced` | quorum, and the claim is recorded as unfounded |
|
|
132
|
+
| `fact` | `not-runnable` | quorum — okstra could not judge it, which refutes nothing |
|
|
133
|
+
| `judgement` | — | quorum |
|
|
134
|
+
| nothing declared | — | blocks on one vote, as before |
|
|
135
|
+
|
|
136
|
+
The last row is deliberate. Verdicts recorded before this field existed are not
|
|
137
|
+
re-judged in hindsight, so adoption only ever relaxes: a claim earns the quorum
|
|
138
|
+
route by declaring itself, never loses a block by staying silent.
|
|
139
|
+
|
|
140
|
+
**Enforced:** `validators/validate-run.py` `_single_vote_block_survives`, with
|
|
141
|
+
the probes in `scripts/okstra_ctl/claim_reproduction.py`.
|
|
142
|
+
|
|
143
|
+
## What the gate asks (BLOCKING)
|
|
144
|
+
|
|
145
|
+
The gate does not ask whether the plan is free of defects. Under adversarial
|
|
146
|
+
reading a plan of any size yields findings every round, so a bar of zero is not
|
|
147
|
+
reachable and a run that aims at it does not end — one task spent five runs and
|
|
148
|
+
its last two went entirely into the plan's account of itself.
|
|
149
|
+
|
|
150
|
+
It asks two things instead, and both must hold:
|
|
151
|
+
|
|
152
|
+
1. **The stage about to start is executable as written.** Its steps, commands,
|
|
153
|
+
exit contract, rollback target and validation signals are settled — §"Gate
|
|
154
|
+
scope" decides which items that covers.
|
|
155
|
+
2. **What was set aside is written down.** Every defect the gate stopped
|
|
156
|
+
blocking on appears in `planBodyVerification.setAside` with its reason —
|
|
157
|
+
`observed` (a frozen stage), `deferred` (a stage not yet reached), or
|
|
158
|
+
`record` (the plan's account of itself).
|
|
159
|
+
|
|
160
|
+
The second condition is what makes the first safe to relax. Knowing a defect is
|
|
161
|
+
the acceptance condition, not removing it; without the register a deferred
|
|
162
|
+
defect and one nobody raised read the same in the report.
|
|
163
|
+
|
|
164
|
+
`okstra plan-items complete-round` writes the register from the same computation
|
|
165
|
+
that produced the gate value. Do not hand-write it — a hand-written register
|
|
166
|
+
drifts from the verdicts it claims to summarise, and nothing else reads it, so
|
|
167
|
+
the drift stays invisible.
|
|
168
|
+
|
|
169
|
+
**Enforced:** `validators/validate-run.py` `_validate_set_aside_register` fails
|
|
170
|
+
a scored round whose declared register does not match the one recomputed from
|
|
171
|
+
the verdicts.
|
|
172
|
+
|
|
173
|
+
## Two blocks — execution and record (BLOCKING)
|
|
174
|
+
|
|
175
|
+
A plan carries two kinds of content, and they answer to different gates.
|
|
176
|
+
|
|
177
|
+
| Block | Items | A defect there |
|
|
178
|
+
|---|---|---|
|
|
179
|
+
| `execution` | steps, options, dependencies, validations, rollbacks, design prep, variation points, the selected direction | blocks the start |
|
|
180
|
+
| `record` | requirement coverage rows, and the approval dispositions and decision references they carry | **does not block.** It is recorded and becomes the next run's input |
|
|
181
|
+
|
|
182
|
+
The record is the plan's account of itself: which brief line each stage answers,
|
|
183
|
+
what the user decided, which clarification a deviation rests on. It has to be
|
|
184
|
+
accurate and it is still verified — but a wrong sentence in it does not make the
|
|
185
|
+
next stage unsafe to start, and treating it as though it did is what kept a plan
|
|
186
|
+
whose executable content had already settled from ever being approved.
|
|
187
|
+
|
|
188
|
+
**A requirement that nothing builds is not a record defect.** A coverage row with
|
|
189
|
+
`status: gap` or a `blocked C-NNN` blocks through `gateBlockedBy: coverage-gap`,
|
|
190
|
+
which reads the rows directly and is untouched by this split. What stops blocking
|
|
191
|
+
is a verdict about the row's *accuracy* — a citation that resolves to the wrong
|
|
192
|
+
stage, a disposition recorded without its confirmation.
|
|
193
|
+
|
|
194
|
+
**Enforced:** `scripts/okstra_ctl/plan_items.py` `_item_block` assigns the block
|
|
195
|
+
at extraction, `okstra plan-items seed` carries it onto the row, and
|
|
196
|
+
`validators/validate-run.py` `_plan_item_gate_class` downgrades a `record`
|
|
197
|
+
blocker to `has-dissent`. `_independent_coverage_blockers` is the separate
|
|
198
|
+
channel that keeps a genuine gap blocking.
|
|
199
|
+
|
|
200
|
+
## Gate scope — the stage about to start (BLOCKING)
|
|
201
|
+
|
|
202
|
+
The plan covers every stage; implementation runs one at a time. An item blocks
|
|
203
|
+
approval only when it has standing over the stage that is about to start:
|
|
204
|
+
|
|
205
|
+
- **`in-scope`** — one of the item's stages is `ready` or `active`, or the item
|
|
206
|
+
belongs to no stage at all. It blocks.
|
|
207
|
+
- **`observed`** — the item's stages are all `done`. It does not block. A frozen
|
|
208
|
+
stage's defect cannot be fixed by planning at all: the Stage Ledger forbids
|
|
209
|
+
editing its commands, so an item that blocks on one blocks forever.
|
|
210
|
+
- **`deferred`** — the item's stages are all still `blocked`. It does not block;
|
|
211
|
+
the stage it judges has not been reached.
|
|
212
|
+
|
|
213
|
+
An item with no `stageScope` is `in-scope` on purpose. `P-Opt-*`, `P-Var-*`,
|
|
214
|
+
`P-Dep-*`, `P-Rb-*` and `P-Dir-1` judge the plan as a whole, and scoping them out
|
|
215
|
+
would stop an unrequested-work verdict from blocking a start. So does a coverage
|
|
216
|
+
or validation row whose plan has not filled `stageRefs` yet — an absent scope is
|
|
217
|
+
read as every stage, which is the safe direction.
|
|
218
|
+
|
|
219
|
+
**Nothing is dropped.** An out-of-scope blocker becomes `has-dissent`, so the
|
|
220
|
+
gate reads `passed-with-dissent` rather than `passed` and the reader can see that
|
|
221
|
+
something is outstanding. `gate.items[]` carries the bucket for each item; a
|
|
222
|
+
scoped-out defect and a real consensus would otherwise look identical in the
|
|
223
|
+
record.
|
|
224
|
+
|
|
225
|
+
**Enforced:** `validators/validate-run.py` `_stage_scope_bucket` and
|
|
226
|
+
`_plan_item_gate_class`, which `_recompute_plan_body_gate` and
|
|
227
|
+
`_gate_summary_item` both call — the scoping cannot apply to the gate value and
|
|
228
|
+
not to the summary the round protocol records.
|
|
229
|
+
|
|
106
230
|
## Plan-body verdict semantics
|
|
107
231
|
|
|
108
232
|
The verdict tokens `AGREE` / `DISAGREE` / `SUPPLEMENT` are reused, but their meaning is plan-specific:
|
|
@@ -187,14 +311,14 @@ Exception for `P-Req-*`: verifiers still MUST NOT re-open the original task brie
|
|
|
187
311
|
|
|
188
312
|
## Adversarial plan-body posture
|
|
189
313
|
|
|
190
|
-
When `config.adversarial == true` (the default for `implementation-planning`; see [convergence](./convergence.md) §"Configuration"), the plan-body round runs with an **adversarial posture**. The classification rules and gate arithmetic in §"Round protocol" are UNCHANGED — `majority-disagree`
|
|
314
|
+
When `config.adversarial == true` (the default for `implementation-planning`; see [convergence](./convergence.md) §"Configuration"), the plan-body round runs with an **adversarial posture**. The classification rules and gate arithmetic in §"Round protocol" are UNCHANGED — `majority-disagree` blocks approval, and that class now includes a blocking-kind minority dissent so a 2-AGREE / 1-DISAGREE on `b` / `c` / `e` is not passed silently. Advisory `dissent-isolated` (`DISAGREE(d)`, `P-Rb-*`) still does not block. Adversarial mode changes only *how each verifier evaluates an item*:
|
|
191
315
|
|
|
192
316
|
- The burden of proof sits on the plan: an item earns `AGREE` only if the verifier actively tried to break it and could not.
|
|
193
317
|
- The verifier MUST open the file paths / symbols / commands the item cites and confirm they exist and are **defined** as written. This is the one allowed widening of the lightweight "judge from internal consistency and stated commands / paths" rule — confirming the existence of cited paths is not "re-analyzing the original requirements". The widening stops at *definition*: a build/test command's **execution success** is out of scope here, because the planning worktree has no dependencies installed (§"Planning-time environment gap"). Confirm the script is declared; do not treat its failure to run as evidence against the plan.
|
|
194
318
|
- If a cited path / command / validation signal cannot be confirmed, the verifier responds `DISAGREE(<kind>)` with the applicable breakage kind (a–f); uncertainty resolves toward DISAGREE, not AGREE.
|
|
195
|
-
- **Single-vote-blocking kinds.** A
|
|
319
|
+
- **Single-vote-blocking kinds.** A reproduced `DISAGREE(a)` (cited path/symbol mismatch) on any item other than a `P-Var-*` one, or a reproduced `DISAGREE(f)` on a `P-Req-*` item, blocks on that one vote even if the rest AGREE. **Enforced:** `validators/validate-run.py` `_single_vote_block_survives`. Kinds `b` / `c` / `e` do not auto-block on one unreproduced vote, but a blocking-kind minority with ≥2 participating votes is still `majority-disagree` and goes to the user — the majority does not silently pass it. **Rollback ordering (`d`) never blocks.** Because `a` is reserved for a concrete contradiction between two spelled-out references, an abbreviated path is raised as `b`, never `a`. **Enforced:** `validators/validate-run.py` `_classify_plan_item_gate`.
|
|
196
320
|
|
|
197
|
-
Plan-body verification stays **lightweight** even under this posture — the `verificationMode = "full-reanalysis"` forcing in [convergence](./convergence.md) §"Adversarial Verification Mode" applies to finding convergence only (see §"Mode constraint"); the adversarial posture here only changes verifier behaviour, not the mode. This raises verification *quality* (active refutation, plan-side burden).
|
|
321
|
+
Plan-body verification stays **lightweight** even under this posture — the `verificationMode = "full-reanalysis"` forcing in [convergence](./convergence.md) §"Adversarial Verification Mode" applies to finding convergence only (see §"Mode constraint"); the adversarial posture here only changes verifier behaviour, not the mode. This raises verification *quality* (active refutation, plan-side burden). A reproduced fact (`a`, or `f` on P-Req) still blocks on one confirmed vote. A blocking-kind minority (`b`/`c`/`e`) with ≥2 participating votes goes to the user rather than passing as `has-dissent`. Rollback ordering (`d`) is advisory and never blocks. A lone surviving `DISAGREE` whose peer returned a non-result does NOT block — a worker failure must not make the gate stricter than a healthy roster would.
|
|
198
322
|
|
|
199
323
|
## Round protocol (single round at default `maxRounds=1`)
|
|
200
324
|
|
|
@@ -222,21 +346,23 @@ CLI-wrapper calls follow the planned execution surface after
|
|
|
222
346
|
consume only `modelExecutionValue`. A missing or invalid invocation contract blocks the
|
|
223
347
|
round before any host or provider process starts.
|
|
224
348
|
|
|
225
|
-
1. Lead runs `okstra plan-items
|
|
349
|
+
1. Lead runs `okstra plan-items prepare --narrative <report-writer-narrative.md> --run-manifest <run-manifest>`, places the fixed output of `okstra plan-items prompt --run-manifest <run-manifest>` verbatim in every verifier prompt, then runs `okstra plan-items validate-prepared --narrative <report-writer-narrative.md> --run-manifest <run-manifest>`. After a self-fix rewrite, pass `--state <plan-body-verification.json>` on prepare and validate-prepared so the dispatch queue is the changed items plus their stage closure, not the full extract. Python resolves the one convergence-owned state path from that run identity. Dispatch only after that exact-match validation succeeds. The prompt is the dispatch queue: `observed` / `deferred` stages are omitted; plan-wide items (`P-Dir-1`, `P-Var-*`, `P-Dep-*`, items with no `stageScope`) stay. **Enforced:** `okstra_ctl.plan_items.dispatch_item_ids` / `reverify_item_ids`.
|
|
226
350
|
|
|
227
|
-
**Then seed the landing table (BLOCKING):** `okstra plan-items seed --narrative <report-writer-narrative.md> --state <plan-body-verification.json>`. `apply-verdicts` in step 8 refuses a verdict whose item has no `planBodyVerification.planItems[]` row. The report writer never owns that state, so the deterministic seed is the only creator of its rows. The seed is idempotent by id and never touches
|
|
351
|
+
**Then seed the landing table (BLOCKING):** `okstra plan-items seed --narrative <report-writer-narrative.md> --state <plan-body-verification.json> --run-manifest <run-manifest.json>`. `apply-verdicts` in step 8 refuses a verdict whose item has no `planBodyVerification.planItems[]` row. The report writer never owns that state, so the deterministic seed is the only creator of its rows. The seed is idempotent by id and never touches existing verdicts, so it is safe to re-run between rounds and after a self-fix re-extraction. It does refresh `contentHash` from the current extract. Skipping it makes step 8 fail with `plan-body state has no row for [...]`.
|
|
352
|
+
|
|
353
|
+
**`--run-manifest` is what scopes the gate to the stage you are starting.** Seed uses it to overlay disk `done` / `active` onto `planBodyVerification.stageLedger`. The current plan's depends-on fills `ready` / `blocked` when no prior plan exists, so a first run does not treat every stage as in-scope. **Enforced:** `okstra_ctl.plan_items.planning_stage_ledger`.
|
|
228
354
|
2. For each analyser worker in the roster (`claude`, `codex`, and `antigravity` if opted in), lead constructs a reverify prompt using the template in §"Plan-body reverify prompt" below.
|
|
229
355
|
3. Dispatch uses the same wrapper infrastructure as finding convergence, so the `--role-slug` is the same canonical `<role>-worker` that convergence uses — not a round-specific slug. Result file path: `runs/<task-type>/worker-results/<role>-worker-plan-verify-r<N>-implementation-planning-<seq>.md` (e.g. `codex-worker-plan-verify-r1-implementation-planning-003.md`). **`<seq>` is the report's sequence** — the one in this run's `final-report-<task-type>-<seq>` filename, NOT the `workerResults` sequence the initial analysis results carry. The two are equal in most runs and diverge in some (`reports: 004` alongside `workerResults: 005` is a real case), and provenance globs on the report's. Picking the other one makes `_validate_plan_body_verdict_provenance` report that no result file exists while the file is sitting in the directory. The `-worker-` token is load-bearing twice over: §"Plan-body reverify prompt" requires the same anchor headers as convergence, whose `**Audit sidecar path:**` is derived by `okstra_ctl.worker_artifact_paths.audit_sidecar_rel()` inserting `-audit-` after that token — a slug without it makes the header underivable and the helper raises. Record each `planItems[].verdicts[].worker` as the same `<role>-worker` string, because provenance compares it to this filename's prefix. **Enforced:** `tests/contract/test_reverify_dispatch_anchors.py` derives the sidecar from the documented name and re-extracts the prefix the provenance resolver uses.
|
|
230
356
|
**Verdict provenance (BLOCKING).** Every verdict recorded in `planItems[].verdicts[]` MUST trace back to a dispatch that actually returned a result file at the path above. The whole gate — classification, self-fix eligibility, promotion, `gateBlockedBy` — is computed from these votes, so an unbacked vote lets the round be skipped while the gate still reads `passed`. **Enforced:** `validators/validate-run.py` `_validate_plan_body_verdict_provenance` fails any `verdicts[].worker` with no matching `<worker>-plan-verify-r<N>-<task-type>-<seq>.md` result file. Recording a `verification-error` for a dispatch that produced no result is the correct way to represent a failed worker — inventing an `AGREE` is a contract violation.
|
|
231
357
|
|
|
232
358
|
4. After all dispatches return, lead aggregates verdicts per `P-*` item across workers and classifies each:
|
|
233
359
|
|
|
234
|
-
**Every item carries at least one verdict row (BLOCKING).** Aggregation covers the
|
|
360
|
+
**Every in-scope item carries at least one verdict row (BLOCKING).** Aggregation covers the dispatch queue, not deferred or observed stages. An in-scope item left with an empty `verdicts[]` is not a weak signal the gate can discount — it classifies `all-non-result`, states as `needs-reverify`, and folds into `passed-with-dissent` next to items two verifiers actually agreed on, so a plan item nobody judged reads as a passing one. This is the shape a self-fix round produces when the planner adds an in-scope item and the targeted round-N queue never picks it up. A worker that returned nothing is a `verification-error` row (step 3), not a missing row; if an in-scope item was never dispatched, dispatch it before scoring the round. **Enforced:** `validators/validate-run.py` `_validate_round_recorded_verdicts` fails any run whose `roundCount` ≥ 1 leaves an in-scope item with no verdict row.
|
|
235
361
|
|
|
236
362
|
- `full-consensus` — all participating analysers `AGREE` (SUPPLEMENT counts as agree on the item itself).
|
|
237
|
-
- `partial-consensus` — majority `AGREE
|
|
238
|
-
- `dissent-isolated` — only one worker `DISAGREE`s, others `AGREE`
|
|
239
|
-
- `majority-disagree` — a *majority* of analysers `DISAGREE` (majority needs ≥2 participating non-error votes; rollback-ordering `DISAGREE(d)` votes are advisory and excluded from the tally), OR any single-vote-blocking kind fires: one `DISAGREE(a)` on any item other than a `P-Var-*` one
|
|
363
|
+
- `partial-consensus` — majority `AGREE` with two or more blocking `DISAGREE`s. On kinds `b` / `c` / `e` this is scored `majority-disagree` and **blocks approval** so the user decides; it is not folded into a passing gate.
|
|
364
|
+
- `dissent-isolated` — only one worker `DISAGREE`s, others `AGREE`. On a blocking kind (`b` / `c` / `e`, and kind `a` on `P-Var-*`) this is scored `majority-disagree` and **blocks approval**. Advisory-only `DISAGREE(d)` and `P-Rb-*` stay recorded dissent and do not block. (Distinct from finding-convergence `worker-unique`, which means the *opposite*: only one worker AGREEs.)
|
|
365
|
+
- `majority-disagree` — a *majority* of analysers `DISAGREE` (majority needs ≥2 participating non-error votes; rollback-ordering `DISAGREE(d)` votes are advisory and excluded from the tally), OR any blocking-kind dissent with ≥2 participating votes (a minority `DISAGREE` is not outvoted), OR any single-vote-blocking kind fires: one reproduced `DISAGREE(a)` on any item other than a `P-Var-*` one, or one reproduced `DISAGREE(f)` on a `P-Req-*` item (see §"Single-vote-blocking kinds"). This classification **blocks approval**.
|
|
240
366
|
- `needs-reverify` — one of two shapes the round could not settle.
|
|
241
367
|
- **An even split on a blocking kind.** The majority test is strict, so a panel splitting evenly (1-AGREE / 1-DISAGREE, 2-2, …) reaches neither `full-consensus` nor `majority-disagree`. Until this shape existed it folded into `has-dissent` and the gate passed: two verifiers read the same plan, disagreed on a defect that is not advisory, and the split was recorded and never acted on. An even panel is not only the two-analyser roster — one `UNVERIFIABLE` or one lost dispatch makes any roster even for that item. Re-dispatch those items and record the votes with `--round 2`; a split that survives that round becomes `majority-disagree` and goes to the user, because nothing further is going to settle it. **The round is not optional**: `needs-reverify` folds into `passed-with-dissent`, so without the re-verification this classification would be a label and nothing else. **Enforced:** `validators/validate-run.py` `_validate_unresolved_tie_was_reverified` fails a gate declared over a tie that was never re-verified, and `_classify_plan_item_gate` promotes a tie carrying a round-2 verdict to `majority-disagree`.
|
|
242
368
|
- **A lone dissent nobody cross-verified** — a single-vote-blocking kind fired but the item has **fewer than 2 participating non-error votes**, i.e. the lone dissent was never cross-verified because its peer returned `verification-error`. A single-vote-blocking kind means "one *confirmed* DISAGREE is enough"; an unconfirmed one is not, and on a `P-Var-*` item none fires at all — its kind `a` never blocks on one vote and takes a majority like `b` / `e`. This does **not** block approval — blocking on it would make a worker failure produce a stricter gate than a healthy roster, the same paradox the ≥2-vote majority rule already rules out. The item is re-dispatched in the next round (step 7); if it survives the round budget it is promoted per step 8 with a Statement that says verification never completed. **Enforced:** `validators/validate-run.py` `_classify_plan_item_gate` returns `needs-reverify` for this shape and `_recompute_plan_body_gate` folds it into `passed-with-dissent`.
|
|
@@ -244,7 +370,7 @@ round before any host or provider process starts.
|
|
|
244
370
|
5. Gate result resolution:
|
|
245
371
|
- any `majority-disagree` item present AND `gating=true` → `blocked-by-disagreement`
|
|
246
372
|
- all dispatches non-result → `aborted-non-result`
|
|
247
|
-
- any
|
|
373
|
+
- any advisory `dissent-isolated` / `needs-reverify` present, no `majority-disagree` → `passed-with-dissent`
|
|
248
374
|
- all items `full-consensus` → `passed`
|
|
249
375
|
|
|
250
376
|
**Score the gate with `okstra plan-verify`, never by hand (BLOCKING).** Once this round's verdicts are in the data.json, lead runs
|
|
@@ -260,27 +386,35 @@ round before any host or provider process starts.
|
|
|
260
386
|
**Record the cause, not just the outcome.** The gate value names the outcome; `planBodyVerification.gateBlockedBy` (array) names every input that blocked it — `majority-disagree`, `coverage-gap`, `non-result`. Two independent inputs can block: a `majority-disagree` plan item, and a Requirement Coverage `gap` / `blocked C-NNN` row (`prompts/profiles/implementation-planning.md` §"Requirement Coverage"). A coverage-only block still renders as `blocked-by-disagreement` because that is the only blocking non-abort value, so **without `gateBlockedBy` the report asserts a worker disagreement that never happened** and the reader hunts for a dissent that does not exist. Leave the array empty for a passing gate. **Enforced:** `validators/validate-run.py` `_validate_gate_blocked_by` cross-checks the declared causes against the recorded verdicts and coverage rows, and fails a passing gate that has a blocking coverage row — the coverage rule was prose-only before.
|
|
261
387
|
|
|
262
388
|
**A coverage row citing this run's own `C-NNN` is not an independent blocker.** When a coverage row's `blocked C-NNN` points at a clarification that step 8 below promoted from a `majority-disagree` item in *this same run*, that blocker is already counted once as the plan item. Counting it again as a coverage gap makes the run block on a clarification it just authored, and the row carries into the next run as a fresh blocker — the Requirement Coverage ↔ Clarification cycle. Such rows are excluded from `coverage-gap`. **Enforced:** `validators/validate-run.py` `_independent_coverage_blockers`.
|
|
263
|
-
6.
|
|
389
|
+
6. `okstra plan-items complete-round --run-manifest <current-run-manifest.json>` derives `planBodyVerification.participatingAnalysers` from the current assigned roster and persisted votes, then atomically records the completed round. The gate arithmetic is unchanged, but a shrunken roster changes what the round can settle: with two participating analysers a 1-AGREE / 1-DISAGREE split is a tie, so it reaches neither consensus nor `majority-disagree` and the item has to go back for a round (see `needs-reverify` above). **Enforced:** `validators/validate-run.py` `_validate_participating_analysers` recomputes `voting` from the recorded verdicts and fails a declared figure the table denies. `validators/validate-run.py` `_detect_uniform_verifier` remains advisory; do not copy its JSON output into state.
|
|
390
|
+
|
|
391
|
+
**Check each verifier's verdict distribution before the next round.** After `apply-verdicts` and before opening another worker batch, run `okstra plan-items next-dispatch --state <plan-body-verification.json> --run-manifest <current-run-manifest.json>`. Python owns that decision. Do not invent a full-roster round from a `needs-reverify` label, from every-item `UNVERIFIABLE`, or from `okstra plan-verify` warnings. `_detect_uniform_verifier` remains advisory; do not copy its JSON output into state.
|
|
264
392
|
|
|
265
|
-
|
|
393
|
+
| `kind` | What the lead does |
|
|
394
|
+
|---|---|
|
|
395
|
+
| `none` | Do not add a worker batch. A round whose only failures are missing-dependency command runs — `UNVERIFIABLE` on a declared `npm` / `pytest` / equivalent, §"Planning-time environment gap" — is this shape. |
|
|
396
|
+
| `worker-correction` | Re-dispatch **only** those workers. Peers are not re-run. The queue does not become a new round. Place the output of `okstra plan-items correction-prompt --worker <id> --run-manifest … --state …` first in that worker's prompt — the environment-exception paragraph is first. A byte-identical re-dispatch reproduces the same failure; a corrected one recovered 37 substantive verdicts from a worker whose first attempt answered `UNVERIFIABLE` to all 80 items. |
|
|
397
|
+
| `queue-reverify` | An unsettled tie on a blocking kind. Re-dispatch those `itemIds` only. |
|
|
266
398
|
|
|
267
|
-
|
|
399
|
+
A referenced **path** that does not exist is still `DISAGREE(b)` / a fact probe, never environment-unverifiable. **Enforced:** `okstra_ctl.plan_items.next_dispatch` / `correction_prompt_text`.
|
|
400
|
+
|
|
401
|
+
The environment exception in §"Planning-time environment gap" covers **running build and test commands only** — whether a referenced path exists, whether a command is declared in `package.json`, and whether the plan is internally consistent are all checkable without it, and a blanket "capability constraints prevent workspace resolution" is not a valid answer to any of them.
|
|
268
402
|
|
|
269
403
|
**How the corrective round is recorded.** The first prompt was dispatched, so it is immutable — `--replace-undispatched` refuses it, correctly. Materialize the correction under a NEW `--invocation-id` and a new prompt path. Before linking its result, retire the first attempt's link: `okstra agent-prompt reject-result --run-manifest <path> --dispatch-id <first dispatch id> --superseded-by <corrective dispatch id> --reason "<what was wrong with the returned result>"`. Without that step the corrective `link-result` fails with `agent result is already linked to another dispatch`, which is how a worker that ran for twenty minutes and wrote a good result ends up unrecordable. Nothing is deleted: the rejected link stays in `agentResultLinks` carrying `supersededBy` and `rejectionReason`, so the ledger shows both attempts and why the second exists.
|
|
270
404
|
|
|
271
|
-
Then
|
|
405
|
+
Then run `okstra plan-items complete-round --state <plan-body-verification.json> --run-manifest <current-run-manifest.json> --round <N>`. Python appends one immutable round history entry, records each verified item's votes, derives the current projection from the actual assigned roster, and stamps `completedAt` after the preceding verification command succeeds. The file accumulates across rounds; it is never truncated to the latest one. Report assembly later projects the completed nested `planBodyVerification` into the final record.
|
|
272
406
|
7. **Self-fix loop (one rewrite, targeting planner-fixable defects).** After round 1, lead may run one report-writer rewrite when at least one `majority-disagree` item has a majority of its `DISAGREE` verdicts at `fixability == planner-fixable`. The targeted re-verification after that rewrite is round 2. After round 2, stop automatic self-fix regardless of outcome. Classify every remaining item as `user-decision`, `noncritical-dissent`, or `correctness-critical`. A second automatic self-fix is a contract violation. The fixed order is initial verification → one planner self-fix → targeted re-verification → user gate.
|
|
273
|
-
- **Group the targets by cause before instructing (BLOCKING).** Blocked items are usually several derivatives of one defect
|
|
407
|
+
- **Group the targets by cause before instructing (BLOCKING).** Blocked items are usually several derivatives of one defect. Lead partitions this round's targets into cause groups and instructs each group as **"remove this cause"**, naming the derivatives it accounts for. The convergence command owns the persisted group and correction fields; the lead does not edit JSON state.
|
|
274
408
|
- lead instructs report-writer to rewrite the items in each cause group (NOT a full draft regeneration; procedure in [report-writer](./report-writer.md) §"Self-fix rewrite").
|
|
275
409
|
- missing or weak `P-Prep-*` contracts are repaired by adding kind-specific inline detail or an AI-prepared PREP item with a concrete proposal. Facts that require user or external authority remain `blocked` and keep their request material; never invent those facts during self-fix.
|
|
276
410
|
- **Drop plan items whose element the round deleted.** A self-fix rewrite may remove a plan element (a validation check, a rollback row). `P-*` ids are positional, so a deletion shifts every later row and silently re-points surviving verdicts at their neighbours — and a verdict recorded against a removed element keeps blocking a gate while being unfindable in the plan, so reading the plan never reveals the cause. After each round, re-extract plan items with `okstra plan-items extract` and re-verify any item whose `subject` no longer matches; never carry the old vote forward across a shift. **Enforced:** `validators/validate-run.py` `_validate_verdicts_match_current_subjects` (re-pointing) and `_validate_plan_item_extraction_completeness` (dangling ids).
|
|
277
411
|
- **Classify each cause group before instructing it (BLOCKING).** A group is either an *authoring* defect — the plan says something wrong, incomplete, or self-contradictory, which self-fix owns — or a *citation* defect, where the plan points at an analysis artifact incorrectly. Only the first is self-fix work. For the second the finding already exists and already went through convergence, so the fix is to re-cite the converged artifact; instructing report-writer to re-derive the fact means the author reads the source material and produces a **finding that never went through convergence**, which the plan then carries as if it had. That is the role boundary the lead contract draws ("keep analysis, execution, verification, and report authoring responsibilities distinct; return defects to the role that owns them"), and report-writer is authoring-only by its own contract. `P-Req-*` items with breakage kind `f` are where this goes wrong most often: the question is usually whether a coverage row points correctly at something already measured, not whether the measurement is right. State the classification in the group's instruction so the author knows which of the two it is being asked to do.
|
|
278
|
-
- **A verdict older than the last self-fix is not a verdict (BLOCKING).** A verdict cast in round 1 judged the text before the only automatic rewrite. Once that rewrite runs,
|
|
279
|
-
-
|
|
280
|
-
-
|
|
412
|
+
- **A verdict older than the last self-fix is not a verdict unless the item's content is unchanged (BLOCKING).** A verdict cast in round 1 judged the text before the only automatic rewrite. Once that rewrite runs, a changed item's judgement is about a plan that no longer exists. `--round <N>` on `apply-verdicts` stamps each row and copies `contentHash` onto `verifiedContentHash`. `validators/validate-run.py` `_validate_verdict_rounds_outlive_self_fix` fails an in-scope item whose verdict round is at or before `selfFixRoundsApplied` **and** whose `contentHash` does not match `verifiedContentHash`. Matching hashes keep the prior verdict — that is what avoids a sweep round over unchanged stages. Deferred and observed items are out of the gate and do not need a post-self-fix verdict. **Enforced:** `_validate_verdict_rounds_outlive_self_fix`.
|
|
413
|
+
- Lead re-runs plan-body verification, then records each worker Markdown result through `okstra plan-items apply-verdicts --state <plan-body-verification.json> --result <worker>=<result.md> --round <N>`. Score the result with `okstra plan-verify --narrative <report-writer-narrative.md> --state <plan-body-verification.json>`, then call `okstra plan-items complete-round --state <plan-body-verification.json> --run-manifest <current-run-manifest.json> --round <N>`. These commands fail on an assigned item the worker left unanswered, on a verdict for an item outside the queue, and on a duplicate worker result.
|
|
414
|
+
- For a self-fix, record the correction through the typed convergence command rather than writing `selfFixNote` or `selfFixGroups` JSON. A resolved item does not create a clarification.
|
|
281
415
|
- **Each round is a worker batch.** Before dispatching round N ≥ 2, reclaim the previous round's completed verifiers exactly as at any other batch boundary ([okstra-lead-contract](./okstra-lead-contract.md) "Run-scoped worker-resource lifecycle") and emit `PROGRESS: phase-batch-cleanup panes=<n>`, then announce the round with `PROGRESS: phase-5.5.9-plan-verify round=<N> items=<count>`. Saying a round will "reuse" the previous verifiers and then dispatching under fresh names leaves every prior round holding its panes — five rounds of that is what exhausts the pane budget and blocks the next dispatch. **Enforced:** `validators/validate_session_conformance.py` `_check_plan_verify_cleanup_checkpoints` requires both lines once the state file records two or more rounds.
|
|
282
|
-
- **Round completion.** A round is complete only after
|
|
283
|
-
- **Loop termination.**
|
|
416
|
+
- **Round completion.** A round is complete only after `okstra plan-verify` exits 0 and `okstra plan-items complete-round` succeeds. A round left with a non-zero exit carries its defect into the next round's inputs. Report assembly and rendering occur only after the convergence state is terminal. **Enforced:** `validators/validate-run.py` `_validate_plan_body_state_rounds` requires one stored round per round number and a corresponding item vote.
|
|
417
|
+
- **Loop termination.** The convergence command owns the round count and stop reason; the lead does not write either JSON field. A user-directed correction does not consume the automatic self-fix limit, and a verification failure after that correction does not restart the automatic loop:
|
|
284
418
|
- `all-resolved` — no planner-fixable `majority-disagree` item remains. Exit.
|
|
285
419
|
- `no-progress` — the round resolved **zero** planner-fixable items relative to the previous round. Exit even with budget left: the same rewrite would repeat. Newly *introduced* defects count against progress, so a rewrite that trades one defect for another stops the loop rather than churning. A round that re-targets only what the previous round left unresolved is this same conclusion reached one dispatch earlier — exit on it under this reason rather than paying for the round that proves it. **Enforced (advisory):** `validators/validate-run.py` `_detect_self_fix_recurrence` warns on that shape and names this stop reason.
|
|
286
420
|
- `max-rounds-reached` — `selfFixRoundsApplied == selfFixMaxRounds`. Exit.
|
|
@@ -293,7 +427,7 @@ round before any host or provider process starts.
|
|
|
293
427
|
- `Statement` summarising the disagreement and the worker breakage `<kind>`
|
|
294
428
|
- `Kind` chosen per the standard policy (usually `decision` for option-level conflicts, `data-point` for path/symbol mismatches)
|
|
295
429
|
- `Blocks=approval`
|
|
296
|
-
- the
|
|
430
|
+
- the report record's `planItems[].clarificationRefs[]` reaching that `C-<N>`. Under contract v3 the report record carries the **plural** field and the v3.0 schema forbids `clarificationId` on a plan item; report assembly derives the refs from the activity ledger's `clarificationRefs[]` + `planItemIds[]`, so record the decision through `okstra approval-decision` rather than writing the link by hand. The lead-owned state file keeps the singular `clarificationId`. `validators/validate-run.py` `_validate_plan_body_clarification_matching` recomputes each item's class and fails when a majority-disagree item reaches no clarification, or reaches one that is missing or not `blocks: approval`.
|
|
297
431
|
- set `approvalContext.classification` to `user-decision` for a majority `needs-user-input` item, `correctness-critical` for `DISAGREE(a)`, `DISAGREE(f)` on `P-Req-*`, or an independent Requirement Coverage blocker, and `noncritical-dissent` for another surviving majority disagreement.
|
|
298
432
|
- record the decision through `okstra approval-decision open`. Each option carries `disposition`, exactly one `reach`, and optional `scopeEffects`. The activity ledger carries affected `planItemIds` and `clarificationRefs`; the approval row never copies those backtrace IDs. `select` is allowed only for `user-decision`, `accept-risk` only for `noncritical-dissent`, and `request-revision` / `reject` for any classification. `correctness-critical` never offers or records `accept-risk`. **Enforced:** `scripts/okstra_ctl/approval_decisions.py` and report assembly.
|
|
299
433
|
- **Self-fix exhaustion is not risk acceptance.** A `noncritical-dissent` item remains blocking until the user explicitly selects `accept-risk`. Record the user's non-empty original text and existing activity check references through `okstra approval-decision resolve`; report assembly derives the report `resolution`.
|
|
@@ -310,6 +444,18 @@ round before any host or provider process starts.
|
|
|
310
444
|
|
|
311
445
|
## `plan-body-verification-<task-type>-<seq>.json` schema
|
|
312
446
|
|
|
447
|
+
**Take the path from the launch prompt, never from this filename (BLOCKING).**
|
|
448
|
+
The run's `## Run Paths` block renders `Plan-body verification state:` with the
|
|
449
|
+
exact path, and `scripts/okstra_ctl/paths.py` is what computed it. Do not build
|
|
450
|
+
the name from the pattern in this heading: a run carries **two seq families** —
|
|
451
|
+
`state` (team-state, convergence, lead-events, this file) and `reports` (the
|
|
452
|
+
final report and its siblings) — and they diverge whenever the two advance at
|
|
453
|
+
different rates. Writing this file under the report's seq puts it where nothing
|
|
454
|
+
looks: `validators/validate_session_conformance.py` resolves it from the
|
|
455
|
+
team-state name, so the round count reads as 0 and
|
|
456
|
+
`verification-round-completed count must match automatic plan-body rounds=0`
|
|
457
|
+
fails a run whose rounds all ran.
|
|
458
|
+
|
|
313
459
|
**Which file is authoritative for what.** Contract v3 keeps both views in one convergence-owned state file before publication:
|
|
314
460
|
|
|
315
461
|
| | records | what a self-fix round does to it |
|
|
@@ -409,9 +555,9 @@ The per-round structures mirror the finding-convergence state artifact ([converg
|
|
|
409
555
|
| `gate.items[].classification` | `planItems[].rounds[].classification` | Condition |
|
|
410
556
|
|---|---|---|
|
|
411
557
|
| `full-consensus` | `full-consensus` | no `DISAGREE` |
|
|
412
|
-
| `has-dissent` | `dissent-isolated` | exactly one `DISAGREE` |
|
|
413
|
-
| `has-dissent` | `partial-consensus` | two or more `DISAGREE
|
|
414
|
-
| `majority-disagree` | `majority-disagree` |
|
|
558
|
+
| `has-dissent` | `dissent-isolated` | exactly one advisory `DISAGREE` (`d` or `P-Rb-*`) |
|
|
559
|
+
| `has-dissent` | `partial-consensus` | two or more advisory `DISAGREE`s |
|
|
560
|
+
| `majority-disagree` | `majority-disagree` | majority, blocking-kind minority, or reproduced single-vote |
|
|
415
561
|
| `needs-reverify` | `needs-reverify` | — |
|
|
416
562
|
| `all-non-result` | `needs-reverify` | no non-error vote at all |
|
|
417
563
|
|
|
@@ -428,9 +574,35 @@ Required prompt anchor headers are identical to finding convergence (see [conver
|
|
|
428
574
|
The [convergence](./convergence.md) §"Required reverify output contract"
|
|
429
575
|
applies unchanged: append it verbatim after the response format below.
|
|
430
576
|
|
|
577
|
+
The prompt-body contract check runs on every `analysis`-audience prompt, not
|
|
578
|
+
just the critic's, so this round needs the same two lines the critic's
|
|
579
|
+
instructions need ([convergence](./convergence.md) §"What the critic
|
|
580
|
+
task-instructions file MUST contain"): a `**Prompt Delivery Mode:**` header and,
|
|
581
|
+
under `## Inputs`, exactly one `- Primary analysis packet:` line whose path ends
|
|
582
|
+
in `analysis-packet.md`. The literal label and the backticks are what the check
|
|
583
|
+
matches — a bare path or a reworded label counts as zero. They are in the
|
|
584
|
+
template below; keep them when you fill it in.
|
|
585
|
+
|
|
586
|
+
Carrying the packet does not license re-analysis. It is there so the verifier
|
|
587
|
+
can resolve a plan item back to the requirement it claims to satisfy; the
|
|
588
|
+
posture in §"Adversarial plan-body posture" still applies, and this round does
|
|
589
|
+
not revisit the requirements themselves.
|
|
590
|
+
|
|
591
|
+
Omitting either line fails `okstra team dispatch --dispatch-kind
|
|
592
|
+
reverify-planbody-r<n>` before any process starts, reported as `<task-type>
|
|
593
|
+
prompt contract: <worker>: exactly one Primary analysis packet path is required
|
|
594
|
+
(found 0)`. Fix the instructions file and re-materialize with
|
|
595
|
+
`--replace-undispatched` rather than editing the published prompt.
|
|
596
|
+
|
|
431
597
|
````
|
|
432
598
|
Perform plan-body verification for <task-key> (round 1).
|
|
433
599
|
|
|
600
|
+
**Prompt Delivery Mode:** eager-include
|
|
601
|
+
|
|
602
|
+
## Inputs
|
|
603
|
+
|
|
604
|
+
- Primary analysis packet: `<path ending in analysis-packet.md>`
|
|
605
|
+
|
|
434
606
|
## Instructions
|
|
435
607
|
|
|
436
608
|
Review the following items extracted from the consolidated implementation plan
|
|
@@ -447,6 +619,13 @@ verdict:
|
|
|
447
619
|
(e) item contradicts the trade-off matrix — **including a `P-Opt-*` option carrying an abstraction (helper module, strategy / factory, indirection layer, interface), a configuration knob every planned call site passes identically, or an optional parameter no planned call site supplies, when no `P-Req-*` requirement-coverage item in this same queue maps to it**: the matrix priced a complexity the option does not buy. Decide this from the queue alone — the requirement-coverage rows are in it, so this is plan-internal consistency, not a re-analysis of the brief. A behavior that already has two implementations is the opposite defect and belongs to `P-Var-*`; do not raise both on one behavior,
|
|
448
620
|
(f) requirement coverage row does not map the stated requirement to a concrete satisfying option / stage / step — citing an existing option counts as concrete even if that option's paths are abbreviated (that is (b) on the option's item, not (f)).
|
|
449
621
|
When you give a DISAGREE, also answer **Fixability** — `planner-fixable` if this defect can be fixed using only the code + this plan draft + the brief, `needs-user-input` if an open user clarification / external information is required.
|
|
622
|
+
|
|
623
|
+
On a `DISAGREE(a)`, or a `DISAGREE(f)` on a `P-Req-*` item, also answer **Claim** — those are the kinds that can block on your vote alone, so state what sort of claim it is and give okstra something to check.
|
|
624
|
+
|
|
625
|
+
- **`fact`** — write a **Probe** okstra can run, as one line of JSON: `{"kind": "path-absent", "path": "src/missing.ts"}`. The kinds are `path-exists`, `path-absent`, `literal-present`, `literal-absent` (a literal inside a file), and `citations-differ` (two spelled-out references that contradict each other, as `left` / `right`). Paths are project-relative.
|
|
626
|
+
- **`judgement`** — no probe. It takes a quorum, like `b` / `c` / `e`.
|
|
627
|
+
|
|
628
|
+
**Do not write a reproduction result.** okstra runs the probe and records the outcome; a result you write is discarded. If your claim cannot be written as one of those probes, it is a `judgement` — that is a classification, not a demotion, and a claim no machine can confirm should never have blocked on one vote.
|
|
450
629
|
- **SUPPLEMENT**: The item is sound but a dependency / edge case / precondition
|
|
451
630
|
is missing.
|
|
452
631
|
- **UNVERIFIABLE**: Capability, credential, network, or service state prevents
|
|
@@ -528,6 +707,8 @@ never recorded, and the round is scored on the items that remain.
|
|
|
528
707
|
### P-Step-3
|
|
529
708
|
**Verdict**: AGREE | DISAGREE(<a|b|c|d|e|f>) | SUPPLEMENT | UNVERIFIABLE
|
|
530
709
|
**Fixability** (only when DISAGREE): planner-fixable | needs-user-input — "planner-fixable if it can be fixed using only the code + this plan + the brief". Fixability is not applicable for `UNVERIFIABLE`.
|
|
710
|
+
**Claim** (only on `DISAGREE(a)`, or `DISAGREE(f)` on a P-Req item): fact | judgement
|
|
711
|
+
**Probe** (only when Claim is fact): <one line of JSON, e.g. {"kind": "path-absent", "path": "src/missing.ts"}>
|
|
531
712
|
**Note**: <the falsification candidate considered and why it is excluded; required even for AGREE>
|
|
532
713
|
**Explanation**: <2-3 sentences>
|
|
533
714
|
|
|
@@ -30,6 +30,25 @@ Resolution `checkRefs` name existing `A-NNN` activity rows. Those activity rows
|
|
|
30
30
|
|
|
31
31
|
## Report-writer dispatch
|
|
32
32
|
|
|
33
|
+
For report contract 3.0, prompt materialization first freezes one report synthesis
|
|
34
|
+
packet. The packet contains the task brief, analysis packet, report template,
|
|
35
|
+
report schema, convergence state, every successful settled worker result recorded
|
|
36
|
+
by the run manifest, accumulated `user-responses/` sidecars, and the current
|
|
37
|
+
session/token/cost accounting snapshot. Each file-backed source carries its owner,
|
|
38
|
+
project-relative path, SHA-256 digest, and value. The report writer receives the
|
|
39
|
+
packet's Markdown reading projection as the only task input instead of an
|
|
40
|
+
independently assembled list of raw paths.
|
|
41
|
+
|
|
42
|
+
The packet's authoring contract names the result path, narrative format, writing
|
|
43
|
+
instructions, runtime-owned content, and validation rules. Missing configured
|
|
44
|
+
paths and missing files are collected across the full source set before dispatch
|
|
45
|
+
and reported together with their owners. Report assembly compares every frozen
|
|
46
|
+
source digest again and reports every changed or missing source in one result.
|
|
47
|
+
This is enforced by
|
|
48
|
+
`initial_prompt_materialization._materialize_report_writer_packet()`,
|
|
49
|
+
`report_synthesis_packet.build_report_synthesis_packet()`, and
|
|
50
|
+
`report_assembly.assemble_report()`.
|
|
51
|
+
|
|
33
52
|
Materialize the duty prompt with `okstra agent-prompt materialize --audience report-writer`. The prompt starts with these anchors in order:
|
|
34
53
|
|
|
35
54
|
1. `**Project Root:**`
|
|
@@ -41,7 +60,7 @@ Materialize the duty prompt with `okstra agent-prompt materialize --audience rep
|
|
|
41
60
|
7. `**Errors log path:**`
|
|
42
61
|
8. `**Errors sidecar path:**`
|
|
43
62
|
|
|
44
|
-
The
|
|
63
|
+
The errors sidecar anchor reserves the runtime-owned write-artifact path used by dispatch validation. No model-authored error JSON file is part of report-writer dispatch; failures use the typed error-log command from the worker error contract.
|
|
45
64
|
|
|
46
65
|
Register the dispatch with `okstra agent-prompt record-dispatch`. When `terminalBackend` is `cmux-pane`, run `okstra team dispatch`; for `runner: cli-wrapper`, run `okstra worker-dispatch --audience report-writer`. Attach the result through `okstra agent-prompt link-result`. The result path is the report narrative Markdown, not the final record.
|
|
47
66
|
|
|
@@ -50,10 +69,10 @@ The pointer record contains the narrative and audit paths. Completion never depe
|
|
|
50
69
|
## Implementation-planning sequence
|
|
51
70
|
|
|
52
71
|
1. Dispatch the report writer and wait for the narrative and pointer.
|
|
53
|
-
2. Parse the narrative and extract the deterministic plan-item queue without publishing `data.json`.
|
|
72
|
+
2. Parse the narrative and extract the deterministic plan-item queue without publishing `data.json`. `okstra plan-items prepare` flips `convergence.planBodyVerification.gating` to `false` when the detector reports `no-design-inputs` and the Stage Map has one row.
|
|
54
73
|
3. Run initial plan-body verification as round 1.
|
|
55
|
-
4. Apply at most one automatic planner self-fix to the narrative.
|
|
56
|
-
5. Run targeted re-verification as round 2 when needed.
|
|
74
|
+
4. Apply at most one automatic planner self-fix to the narrative. Skip this step when `gating` is `false`.
|
|
75
|
+
5. Run targeted re-verification as round 2 when needed. Skip this step when `gating` is `false`.
|
|
57
76
|
6. Persist the completed `planBodyVerification` value in convergence state.
|
|
58
77
|
7. Complete the design-surface detector snapshot.
|
|
59
78
|
8. Run Phase 7 report assembly.
|