okstra 0.189.3 → 0.190.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +1 -1
- package/dist/cli-registry.mjs +9 -0
- package/dist/cli-registry.mjs.map +1 -1
- package/dist/commands/chat/chat.mjs +54 -4
- package/dist/commands/chat/chat.mjs.map +1 -1
- package/dist/commands/lifecycle/install.mjs +12 -5
- package/dist/commands/lifecycle/install.mjs.map +1 -1
- package/dist/commands/lifecycle/uninstall.mjs +2 -1
- package/dist/commands/lifecycle/uninstall.mjs.map +1 -1
- package/dist/lib/host-config.d.mts +5 -3
- package/dist/lib/host-config.mjs +89 -31
- package/dist/lib/host-config.mjs.map +1 -1
- package/docs/architecture/storage-model.md +4 -1
- package/docs/architecture.md +3 -3
- package/docs/cli.md +11 -9
- package/docs/for-ai/skills/okstra-brief-gen.md +1 -1
- package/docs/for-ai/skills/okstra-chat.md +3 -3
- package/docs/for-ai/skills/okstra-inspect.md +11 -1
- package/docs/for-ai/skills/okstra-rollup.md +1 -0
- package/docs/for-ai/skills/okstra-run.md +5 -4
- package/docs/for-ai/skills/okstra-user-response.md +1 -1
- package/docs/project-structure-overview.md +17 -5
- package/docs/task-process/README.md +1 -1
- package/docs/task-process/error-analysis.md +2 -2
- package/docs/task-process/final-verification.md +1 -1
- package/docs/task-process/implementation.md +1 -1
- package/docs/task-process/release-handoff.md +8 -6
- package/docs/task-process/requirements-discovery.md +1 -1
- package/package.json +1 -1
- package/runtime/BUILD.json +2 -2
- package/runtime/agents/workers/claude-worker.md +4 -0
- package/runtime/bin/okstra-report-translate.py +10 -4
- package/runtime/prompts/launch.template.md +11 -5
- package/runtime/prompts/lead/adapters/cmux.md +4 -7
- package/runtime/prompts/lead/context-loader.md +2 -0
- package/runtime/prompts/lead/convergence.md +20 -83
- package/runtime/prompts/lead/okstra-lead-contract.md +24 -6
- package/runtime/prompts/lead/plan-body-verification.md +10 -3
- package/runtime/prompts/lead/report-writer.md +15 -5
- package/runtime/prompts/lead/team-contract.md +15 -9
- package/runtime/prompts/profiles/_common-contract.md +2 -1
- package/runtime/prompts/profiles/_implementation-deliverable.md +2 -0
- package/runtime/prompts/profiles/_implementation-diff-review.md +4 -0
- package/runtime/prompts/profiles/_implementation-executor.md +9 -5
- package/runtime/prompts/profiles/_implementation-self-check.md +2 -0
- package/runtime/prompts/profiles/_implementation-verifier.md +3 -4
- package/runtime/prompts/profiles/_stage-discipline.md +12 -5
- package/runtime/prompts/profiles/final-verification.md +1 -1
- package/runtime/prompts/profiles/implementation-planning.md +1 -1
- package/runtime/prompts/profiles/implementation.md +2 -0
- package/runtime/prompts/profiles/release-handoff.md +4 -3
- package/runtime/prompts/wizard/prompts.ko.json +2 -1
- package/runtime/python/okstra_ctl/adapters/dispatch/cmux.py +6 -0
- package/runtime/python/okstra_ctl/adapters/hosts/antigravity/adapter.py +1 -0
- package/runtime/python/okstra_ctl/adapters/hosts/antigravity/relay.md +2 -2
- package/runtime/python/okstra_ctl/adapters/hosts/capability_adapter.py +9 -0
- package/runtime/python/okstra_ctl/adapters/hosts/claude-code/adapter.py +1 -0
- package/runtime/python/okstra_ctl/adapters/hosts/claude-code/relay.md +3 -3
- package/runtime/python/okstra_ctl/adapters/hosts/codex/adapter.py +1 -0
- package/runtime/python/okstra_ctl/adapters/hosts/codex/relay.md +17 -3
- package/runtime/python/okstra_ctl/adapters/hosts/external/adapter.py +1 -0
- package/runtime/python/okstra_ctl/adapters/hosts/external/relay.md +3 -3
- package/runtime/python/okstra_ctl/adapters/hosts/grok/adapter.py +1 -0
- package/runtime/python/okstra_ctl/adapters/hosts/grok/relay.md +2 -2
- package/runtime/python/okstra_ctl/adapters/hosts/kimi/adapter.py +1 -0
- package/runtime/python/okstra_ctl/adapters/hosts/kimi/relay.md +2 -2
- package/runtime/python/okstra_ctl/adapters/providers/codex/adapter.py +3 -3
- package/runtime/python/okstra_ctl/agent/invocation.py +107 -5
- package/runtime/python/okstra_ctl/agent/prompt_cli/cli.py +16 -0
- package/runtime/python/okstra_ctl/agent/prompt_cli/dynamic_verifier.py +3 -2
- package/runtime/python/okstra_ctl/agent/prompt_cli/jobs.py +87 -0
- package/runtime/python/okstra_ctl/agent/prompt_cli/materialize.py +72 -3
- package/runtime/python/okstra_ctl/analysis_packet.py +47 -6
- package/runtime/python/okstra_ctl/cmux.py +34 -13
- package/runtime/python/okstra_ctl/consumers.py +30 -7
- package/runtime/python/okstra_ctl/convergence.py +110 -21
- package/runtime/python/okstra_ctl/convergence_provenance.py +83 -0
- package/runtime/python/okstra_ctl/dispatch_core.py +127 -5
- package/runtime/python/okstra_ctl/dispatch_state.py +17 -9
- package/runtime/python/okstra_ctl/domain/host.py +5 -0
- package/runtime/python/okstra_ctl/domain/wizard/interaction.py +1 -0
- package/runtime/python/okstra_ctl/error_log_write.py +44 -4
- package/runtime/python/okstra_ctl/execution_mutation_audit.py +84 -7
- package/runtime/python/okstra_ctl/fixed_text.py +25 -7
- package/runtime/python/okstra_ctl/group_context.py +475 -10
- package/runtime/python/okstra_ctl/handoff.py +109 -32
- package/runtime/python/okstra_ctl/handoff_verification.py +80 -0
- package/runtime/python/okstra_ctl/initial_prompt_materialization.py +52 -10
- package/runtime/python/okstra_ctl/lead_progress.py +291 -0
- package/runtime/python/okstra_ctl/model_io/renderers.py +60 -1
- package/runtime/python/okstra_ctl/model_io_cli.py +6 -1
- package/runtime/python/okstra_ctl/ports/worker_dispatch.py +4 -0
- package/runtime/python/okstra_ctl/recap.py +179 -17
- package/runtime/python/okstra_ctl/reconcile.py +55 -4
- package/runtime/python/okstra_ctl/render.py +34 -8
- package/runtime/python/okstra_ctl/report_assembly.py +17 -10
- package/runtime/python/okstra_ctl/report_finalize.py +204 -11
- package/runtime/python/okstra_ctl/report_html/common.py +241 -113
- package/runtime/python/okstra_ctl/report_html/context_links.py +121 -0
- package/runtime/python/okstra_ctl/report_html/models.py +11 -5
- package/runtime/python/okstra_ctl/report_html/render.py +69 -17
- package/runtime/python/okstra_ctl/report_html/view_models/change_impact_analysis.py +8 -1
- package/runtime/python/okstra_ctl/report_html/view_models/error_analysis.py +8 -1
- package/runtime/python/okstra_ctl/report_html/view_models/feature_analysis.py +11 -1
- package/runtime/python/okstra_ctl/report_html/view_models/final_verification.py +13 -1
- package/runtime/python/okstra_ctl/report_html/view_models/implementation.py +9 -1
- package/runtime/python/okstra_ctl/report_html/view_models/implementation_option_selection.py +29 -2
- package/runtime/python/okstra_ctl/report_html/view_models/implementation_planning.py +14 -9
- package/runtime/python/okstra_ctl/report_html/view_models/improvement_discovery.py +8 -1
- package/runtime/python/okstra_ctl/report_html/view_models/project_analysis.py +15 -1
- package/runtime/python/okstra_ctl/report_html/view_models/release_handoff.py +7 -1
- package/runtime/python/okstra_ctl/report_html/view_models/requirements_discovery.py +9 -1
- package/runtime/python/okstra_ctl/report_translation.py +58 -1
- package/runtime/python/okstra_ctl/run.py +54 -14
- package/runtime/python/okstra_ctl/stage_targets.py +66 -0
- package/runtime/python/okstra_ctl/verdict_blocks.py +39 -12
- package/runtime/python/okstra_ctl/wizard/__init__.py +148 -0
- package/runtime/python/okstra_ctl/wizard/__main__.py +13 -0
- package/runtime/python/okstra_ctl/wizard/cli.py +176 -0
- package/runtime/python/okstra_ctl/wizard/confirmation.py +331 -0
- package/runtime/python/okstra_ctl/wizard/engine.py +363 -0
- package/runtime/python/okstra_ctl/wizard/ids.py +411 -0
- package/runtime/python/okstra_ctl/wizard/prompts.py +202 -0
- package/runtime/python/okstra_ctl/wizard/registry.py +836 -0
- package/runtime/python/okstra_ctl/wizard/render.py +189 -0
- package/runtime/python/okstra_ctl/wizard/roles.py +737 -0
- package/runtime/python/okstra_ctl/wizard/sources.py +703 -0
- package/runtime/python/okstra_ctl/wizard/state.py +689 -0
- package/runtime/python/okstra_ctl/wizard/statefile.py +145 -0
- package/runtime/python/okstra_ctl/wizard/steps_analysis.py +558 -0
- package/runtime/python/okstra_ctl/wizard/steps_identity.py +828 -0
- package/runtime/python/okstra_ctl/wizard/steps_options.py +340 -0
- package/runtime/python/okstra_ctl/wizard/steps_plan.py +959 -0
- package/runtime/python/okstra_ctl/wizard/steps_roles.py +672 -0
- package/runtime/python/okstra_ctl/worker_prompt_contract.py +43 -3
- package/runtime/python/okstra_ctl/worker_prompt_headers.py +14 -0
- package/runtime/python/okstra_ctl/workflow.py +2 -2
- package/runtime/python/okstra_ctl/worktree/__init__.py +2 -0
- package/runtime/python/okstra_ctl/worktree/cleanliness.py +7 -4
- package/runtime/python/okstra_ctl/worktree/git_ops.py +7 -0
- package/runtime/python/okstra_ctl/worktree/naming.py +15 -6
- package/runtime/python/okstra_ctl/worktree/provision.py +2 -1
- package/runtime/python/okstra_ctl/worktree_registry.py +27 -0
- package/runtime/python/okstra_ctl/wrapper_status.py +4 -0
- package/runtime/schemas/final-report-v2.0.schema.json +4 -0
- package/runtime/schemas/final-report-v3.0.schema.json +4 -0
- package/runtime/skills/okstra-brief-gen/SKILL.md +22 -3
- package/runtime/skills/okstra-chat/SKILL.md +8 -6
- package/runtime/skills/okstra-inspect/SKILL.md +2 -2
- package/runtime/skills/okstra-inspect/facets/recap.md +52 -4
- package/runtime/skills/okstra-rollup/SKILL.md +1 -1
- package/runtime/skills/okstra-run/SKILL.md +17 -7
- package/runtime/skills/okstra-schedule-gen/SKILL.md +2 -0
- package/runtime/skills/okstra-user-response/SKILL.md +1 -1
- package/runtime/templates/reports/brief.template.md +15 -7
- package/runtime/templates/reports/final-verification-input.template.md +2 -0
- package/runtime/templates/reports/group-context.template.md +7 -0
- package/runtime/templates/reports/html/base.template.html +13 -3
- package/runtime/templates/reports/html/i18n/en.json +43 -5
- package/runtime/templates/reports/html/i18n/ko.json +43 -5
- package/runtime/templates/reports/html/macros/forms.html +4 -4
- package/runtime/templates/reports/html/tasks/final-verification.template.html +11 -1
- package/runtime/templates/reports/html/tasks/implementation-option-selection.template.html +16 -15
- package/runtime/templates/reports/html/tasks/implementation-planning.template.html +5 -5
- package/runtime/templates/reports/html/tasks/release-handoff.template.html +2 -2
- package/runtime/templates/reports/html/tasks/requirements-discovery.template.html +1 -1
- package/runtime/templates/reports/implementation-planning-input.template.md +3 -0
- package/runtime/templates/reports/improvement-discovery-input.template.md +3 -0
- package/runtime/templates/reports/release-handoff-input.template.md +5 -3
- package/runtime/templates/reports/schedule.template.md +13 -4
- package/runtime/templates/reports/task-brief.template.md +1 -1
- package/runtime/templates/reverify-output-contract.md +5 -0
- package/runtime/templates/worker-error-contract.md +2 -0
- package/runtime/templates/worker-prompt-preamble.md +1 -1
- package/runtime/validators/validate-report-views.py +30 -1
- package/runtime/validators/validate-run.py +194 -25
- package/runtime/validators/validate_session_conformance.py +102 -2
- package/runtime/python/okstra_ctl/wizard.py +0 -7168
|
@@ -111,10 +111,10 @@ convergence-<task-type>-<seq>.json
|
|
|
111
111
|
Follow this protocol exactly:
|
|
112
112
|
|
|
113
113
|
0. Version-selected schemas describe what the reducer reads: `schemas/convergence-groups-v1.0.schema.json` accepts only legacy groups, `schemas/convergence-groups-v2.0.schema.json` accepts only explicit v2 execution identity, and `schemas/convergence-round-results-v1.0.schema.json` feeds step 4's `apply-round --results`. `schemas/convergence-critic-results-v1.0.schema.json` is a fourth shape but **not** a reducer input — it describes the critic worker's own result document. Step 6's `apply-critic-gaps --results` takes the coverage batch you assemble from those candidates plus each analyser's vote (`{schemaVersion, taskKey, mode, provider, modelExecutionValue, dispatches[], gaps[]}`, spelled out in §"Coverage critic pass" §"State"); feeding the critic document straight in is rejected, by design. `okstra convergence example --kind <groups|round-results|critic-results>` prints a deterministic valid v1 instance of each and writes only JSON to stdout.
|
|
114
|
-
1. Run `okstra convergence seed --groups <groups> --run-manifest <current-run-manifest> --work-state <work> --final-state <final> --migration-dir <state/migrations>` for a v2 worker roster. `--run-manifest` must be the exact current manifest named by the groups document's `runManifestPath`; a previous run from the same task is not interchangeable. Omit the flag for a legacy v1 roster. A `reuse-final` action means validate the existing final and continue to Phase 6. `create-work`, `resume-work`, and `restart-round0` continue with planning.
|
|
114
|
+
1. Run `okstra convergence seed --groups <groups> --run-manifest <current-run-manifest> --work-state <work> --final-state <final> --migration-dir <state/migrations>` for a v2 worker roster. `--run-manifest` must be the exact current manifest named by the groups document's `runManifestPath`; a previous run from the same task is not interchangeable. `seed` refuses a grouping that cites a worker whose initial attempt the mutation audit discarded (`contract-failed-unattributed`, `mutation-present-unresolved`) — the result file is still on disk, but the ledger says it was never accepted. Re-dispatch that worker as a new invocation and cite the new result, or leave it out of the grouping. **Enforced:** `discarded_worker_errors` in `scripts/okstra_ctl/convergence_provenance.py`, called by `_seed`. Omit the flag for a legacy v1 roster. A `reuse-final` action means validate the existing final and continue to Phase 6. `create-work`, `resume-work`, and `restart-round0` continue with planning.
|
|
115
115
|
2. Run `okstra convergence plan-round --work-state <work> --plan <round-plan>`. This is read-only with respect to the working state.
|
|
116
116
|
3. Every plan carries `dispatchable`: `true` on an `action: "dispatch"` plan, `false` on an `action: "finalize"` plan. When the plan action is `dispatch`, create exactly one reverify prompt for each `dispatches[]` row and dispatch it through the selected runtime adapter. Its findings are exactly that row's `findingIds`.
|
|
117
|
-
4.
|
|
117
|
+
4. Run `okstra convergence collect-results --plan <round-plan> --mode <adversarial|collaborative> --result <worker>=<path>… --run-manifest <current-run-manifest> --output convergence-round-<N>-results-<task-type>-<seq>.json`. Each planned worker's terminal outcome and `durationMs` come from the manifest's attempt ledger, not from the lead: an attempt still `started` refuses the collect — run `okstra team await` first so the attempt closes — and a worker whose attempt did not close `ok` cannot contribute a vote even if its result file exists. **Enforced:** `_dispatch_rows_from_manifest` in `scripts/okstra_ctl/convergence.py`. Then run `okstra convergence apply-round --work-state <work> --plan <round-plan> --results <round-results>`.
|
|
118
118
|
5. Repeat `plan-round` and `apply-round` until the plan action is `finalize`. A `finalize` plan is **never dispatched**: it carries `dispatchable: false`, an empty `dispatches[]`, and a `note` naming the next command. Its `round` is only the `<N>` in `convergence-round-<N>-plan-<task-type>-<seq>.json` — not a round to run — so do not open reverify workers for it. **Enforced:** `okstra convergence apply-round` refuses a non-dispatch plan and returns the gate's closing reason with the remedy (`scripts/okstra_ctl/convergence_engine.py` `apply_round_results`).
|
|
119
119
|
6. When the coverage critic is enabled, convert its verification batch to the canonical `dispatches[]` / `gaps[]` result shape and run `okstra convergence apply-critic-gaps --work-state <work> --results <critic-results>` exactly once. The reducer, not the lead, merges verified gaps.
|
|
120
120
|
7. Run `okstra convergence finalize --work-state <work> --output <final>`.
|
|
@@ -128,6 +128,8 @@ After `seed`, the lead MUST NOT hand-edit queue IDs, classifications, `roundHist
|
|
|
128
128
|
|
|
129
129
|
A valid historical final schema v1.0, v1.1, or v1.2 is reused unchanged under `reuse-final`; it remains read-only and is not upgraded in place. A malformed, unreadable, or partial legacy final is archived byte-for-byte under `state/migrations/` and restarted from Round 0 because its queue provenance cannot be reconstructed. A matching valid working state resumes. A malformed new-engine working state fails closed and names the explicit `--restart-from-round0` recovery flag; when supplied, every invalid existing state is archived before replacement. The final path can be replaced only when its recorded migration archive still matches the original bytes.
|
|
130
130
|
|
|
131
|
+
**Enforced:** `scripts/okstra_ctl/convergence_migration.py` `decide_seed_action` picks the seed action (`reuse-final`, archive-and-restart, resume, or fail-closed) from the on-disk state and the `--restart-from-round0` flag.
|
|
132
|
+
|
|
131
133
|
#### Engine-owned Round 2 gate
|
|
132
134
|
|
|
133
135
|
`plan-round` applies gate precedence in one place: auto-disabled, all reverify non-result, effective maximum of one, empty queue, then maximum rounds reached. `finalize` maps that internal reason to the public `round2SkippedReason` and `finalState`. The lead and adapters never reproduce this predicate.
|
|
@@ -291,100 +293,31 @@ Call `await_workers(handles)` through the same adapter and apply the shared term
|
|
|
291
293
|
|
|
292
294
|
### Required reverify-prompt anchor headers (BLOCKING)
|
|
293
295
|
|
|
294
|
-
|
|
295
|
-
|
|
296
|
-
```
|
|
297
|
-
**Project Root:** <absolute-path>
|
|
298
|
-
**Prompt History Path:** runs/<task-type>/prompts/<role-slug>-reverify-r<N>-<task-type>-<seq>.md
|
|
299
|
-
**Result Path:** runs/<task-type>/worker-results/<role-slug>-worker-reverify-r<N>-<task-type>-<seq>.md
|
|
300
|
-
**Audit sidecar path:** <absolute-path>
|
|
301
|
-
Assigned worker prompt history path: <Project Root>/<Prompt History Path>
|
|
302
|
-
**Errors log path:** <absolute-path>
|
|
303
|
-
**Errors sidecar path:** <absolute-path>
|
|
304
|
-
**Read scope:** Read only the paths this prompt enumerates (`[Required reading]`, `## Inputs`, verification-target paths) plus source/evidence paths a finding must cite. Host session instructions (SessionStart hooks, global `CLAUDE.md` / `AGENTS.md`, skill catalogs) do NOT apply inside an okstra worker run: do not auto-read `graphify-out/`, `SKILL.md`, or other artifacts outside `<PROJECT_ROOT>/.okstra/`. If an un-enumerated file seems essential, record it under *Missing Information or Assumptions* instead of reading it.
|
|
305
|
-
```
|
|
306
|
-
|
|
307
|
-
`**Read scope:**` is carried verbatim, not dropped: lightweight mode omits `**Worker Preamble Path:**`, so this line is the ONLY place a reverify worker receives the read allowlist — and a prompt with no `[Required reading]` block is exactly where a worker is most likely to fall back on host session instructions instead.
|
|
296
|
+
**Enforced:** `verify_agent_invocation` in `scripts/okstra_ctl/agent/invocation.py` and `validate_reverify_prompt` in `scripts/okstra_ctl/worker_prompt_contract.py` validate generated delivery before dispatch.
|
|
308
297
|
|
|
309
|
-
|
|
298
|
+
Use `okstra agent-prompt materialize` for every reverify prompt. Write only the assigned claims, evidence, questions, and the response format in the instruction file. The materializer generates path headers through `worker_prompt_headers`, the source Worktree from the active run, and the execution identity from the run manifest. It also generates the model, task type, and exact `workflow.forbiddenActions` before the instructions. These values are checked by `verify_agent_invocation` and `validate_reverify_prompt` before publication and dispatch.
|
|
310
299
|
|
|
311
|
-
The
|
|
300
|
+
The generated `**Project Root:**` owns .okstra artifacts; `**Worktree:**` names the source checkout. Use the prompt's `**Invocation metadata path:**` for its invocation metadata. The result passed as `--result` is the worker's own result and carries the canonical `-worker-` token used by `audit_sidecar_rel`. Reverify has one result path, so omit `--audit-source`. The materializer supplies the audit, errors, read scope, and provider-specific plain-file write instructions.
|
|
312
301
|
|
|
313
|
-
The
|
|
302
|
+
The generated `**Errors log path:**` and `**Errors sidecar path:**` headers use the run's error wiring. Report observed failures with `okstra error-log append-observed` as specified in [team-contract](./team-contract.md). The errors sidecar is runtime-owned; no model-authored error JSON file is part of reverify.
|
|
314
303
|
|
|
315
|
-
|
|
304
|
+
An older instruction file may already contain the fixed boundary. Matching values are retained once; conflicting model, task type, or forbidden-actions values are reported with the expected value and a materialization remedy. Do not edit a dispatched prompt. Use a fresh invocation ID and prompt path for a dynamic reverify correction; undispatched initial prompts can be regenerated in place by the existing recovery path.
|
|
316
305
|
|
|
317
|
-
|
|
306
|
+
After materialization, generate the batch file from the returned metadata paths:
|
|
318
307
|
|
|
319
|
-
|
|
320
|
-
value from `okstra_ctl.worker_prompt_headers` immediately after
|
|
321
|
-
`**Read scope:**` and before the prompt body. Copy it verbatim from that role's
|
|
322
|
-
persisted initial Phase 4 prompt; do not paraphrase or reconstruct it. If the
|
|
323
|
-
persisted initial prompt does not contain that generated header, abort the
|
|
324
|
-
reverify dispatch and record a `contract-violation` event instead of
|
|
325
|
-
dispatching without the plain-file safeguard. This provider-specific header is
|
|
326
|
-
outside the common 8-header count above. Other providers do not receive it.
|
|
308
|
+
Name the reverify result `<role-slug>-worker-reverify-r<N>-<task-type>-<seq>.md` under the run's authorized worker-results directory. Its generated audit path is checked before dispatch.
|
|
327
309
|
|
|
328
|
-
|
|
329
|
-
|
|
330
|
-
The task-instructions file the lead writes MUST open with this block, before any
|
|
331
|
-
`##` heading of its own:
|
|
332
|
-
|
|
333
|
-
```markdown
|
|
334
|
-
**Model:** <role>, <modelExecutionValue>
|
|
335
|
-
**Task Type:** <task-manifest taskType>
|
|
336
|
-
**Forbidden actions:**
|
|
337
|
-
<active-run-context workflow.forbiddenActions, verbatim>
|
|
310
|
+
```bash
|
|
311
|
+
okstra agent-prompt jobs --project-root <root> --run-manifest <manifest> --dispatch-kind reverify-r<N> --metadata <first-meta.json> --metadata <second-meta.json> --out <run-state>/reverify-jobs-r<N>.json
|
|
338
312
|
```
|
|
339
313
|
|
|
340
|
-
|
|
341
|
-
heading — not even the prompt's own opening sentence. The block's value is read
|
|
342
|
-
from the line after `**Forbidden actions:**` up to the next `## ` heading or
|
|
343
|
-
`**X:**` header, so any line placed in that gap is folded into the value and the
|
|
344
|
-
exact-match check fails while the text is plainly correct. This is why each
|
|
345
|
-
prompt-body example below opens at `## Instructions` and states its round line
|
|
346
|
-
inside that section. **Enforced:** `_section_values` in
|
|
347
|
-
`scripts/okstra_ctl/worker_prompt_contract.py`.
|
|
348
|
-
|
|
349
|
-
This is the same placement `okstra_ctl.worker_prompt_body` uses for an initial
|
|
350
|
-
Phase 4 prompt, and it is where the checks look: `validate_reverify_prompt()`
|
|
351
|
-
reads the region after `## Task Instructions`, so a `**Model:**` line left in
|
|
352
|
-
the anchors is invisible to it, and a phase boundary written below the file's
|
|
353
|
-
first heading fails. Both rules are about this file — the composer's own
|
|
354
|
-
`## Duty Contract` heading sits above everything here and is not what they
|
|
355
|
-
measure against.
|
|
356
|
-
|
|
357
|
-
Do not summarize, shorten, or reconstruct the forbidden-actions text. The
|
|
358
|
-
selected adapter validates the task type and exact block through
|
|
359
|
-
`okstra_ctl.worker_prompt_contract.validate_reverify_prompt()` before starting
|
|
360
|
-
the wrapper. `validators/validate-run.py` separately fails the run when the
|
|
361
|
-
run-level error log records a phase-boundary `contract-violation`; a correct
|
|
362
|
-
finding does not make evidence obtained across the phase boundary admissible.
|
|
363
|
-
|
|
364
|
-
`<modelExecutionValue>` MUST be resolved from one of these canonical sources, in priority order:
|
|
365
|
-
|
|
366
|
-
1. `task-manifest.json` → `resultContract.requiredWorkerRoles[].modelExecutionValue` for the receiving role
|
|
367
|
-
2. `team-state-<task-type>-<seq>.json` → `workers[].usage.cliModel` for that role (initial run's actual execution value)
|
|
368
|
-
3. The `**Model:**` line of the initial Phase 4 prompt for that role (read from its persisted prompt-history file)
|
|
369
|
-
|
|
370
|
-
If none of the three is available, **abort the reverify dispatch for that role** and record a `contract-violation` event via `okstra error-log append-observed`. Do NOT guess or fall back to a runtime default. The selected adapter receives the exact value through the assignment and owns how it reaches the worker runtime. The current Codex catalog default is `gpt-5.6-sol`, but that reference value never replaces the manifest assignment.
|
|
371
|
-
|
|
372
|
-
**Enforced:** `okstra_ctl.dispatch_state` hands the job's own `modelExecutionValue` to `validate_reverify_prompt(expected_model=...)`, so a `**Model:**` header that disagrees with the model this dispatch will run fails **before** launch. Carrying a stale value from an earlier round's prompt — the round-1 model into a round-2 header — is the shape this catches; without it the mismatch surfaces as a provider 400 that reads as a worker fault. `tests/contract/test_reverify_dispatch_anchors.py` pins both the check and the dispatch-side hand-off.
|
|
314
|
+
Pass only the current engine-planned batch. The generator reads the actual result headers and canonical role execution, verifies all inputs, then publishes the file. Dispatch the emitted file with the selected adapter's `--jobs-file`; do not copy identity, path, or digest fields by hand. Existing v1 jobs-file consumers remain available.
|
|
373
315
|
|
|
374
316
|
### Required reverify output contract (BLOCKING)
|
|
375
317
|
|
|
376
|
-
|
|
377
|
-
MUST append this block verbatim after its variant-specific response format:
|
|
378
|
-
|
|
379
|
-
**Enforced:** `okstra_ctl.worker_prompt_contract.validate_reverify_prompt` runs at dispatch (`dispatch_state.py`) and rejects a prompt without this block, or with the heading but a missing clause. It matches the three clauses by their distinctive tokens rather than byte-for-byte — a hand-copied block's whitespace must not decide whether a round runs.
|
|
318
|
+
**Enforced:** `validate_reverify_prompt` and `_validate_output_contract_block` in `scripts/okstra_ctl/worker_prompt_contract.py` validate the three persistence clauses before dispatch.
|
|
380
319
|
|
|
381
|
-
|
|
382
|
-
## Output Contract
|
|
383
|
-
|
|
384
|
-
- Write the complete response to the exact `**Result Path:**` before returning.
|
|
385
|
-
- Write the reverify-session audit to `**Audit sidecar path:**` before returning; list the evidence paths actually inspected, or state that the prompt required no external evidence reads.
|
|
386
|
-
- Do not return only an inline response. Inline text is diagnostic output, not the persisted worker result.
|
|
387
|
-
```
|
|
320
|
+
The materializer appends [the canonical output contract](../../templates/reverify-output-contract.md) when the instruction has none. It requires a persisted Result Path and Audit sidecar path before returning and rejects an inline-only response. An existing output block is checked for the same three clauses by `validate_reverify_prompt`. Plan-body verification retains its own response format; its instructions can reference the same output contract.
|
|
388
321
|
|
|
389
322
|
### Reverify prompt: required-reading suppression
|
|
390
323
|
|
|
@@ -506,6 +439,8 @@ needed for a finding, you MUST answer UNVERIFIABLE with a non-empty explanation.
|
|
|
506
439
|
You MUST NOT answer DISAGREE merely because access or reproduction failed;
|
|
507
440
|
UNVERIFIABLE is persisted as `verification-error` and excluded from consensus.
|
|
508
441
|
|
|
442
|
+
**Enforced:** `scripts/okstra_ctl/convergence_engine.py` `_VERDICTS` accepts the verdict as `verification-error`, and `classify_collaborative_round` drops those votes before counting consensus.
|
|
443
|
+
|
|
509
444
|
## Task bundle paths
|
|
510
445
|
- Instruction set: <instruction-set path>
|
|
511
446
|
- Task brief: <task-brief path>
|
|
@@ -756,6 +691,8 @@ Each candidate blocker is verified by the Phase 4 analysers — all of them, on
|
|
|
756
691
|
- **Confirmed** (an analyser reproduces it or cites supporting evidence) → promote to a `## 5.8 Acceptance Blockers` row (keep severity + recommended follow-up phase).
|
|
757
692
|
- **Not confirmed** (cannot reproduce, or evidence is weak) → **downgrade to a Residual Risk row — never drop it.** Record the escalation trigger so the user can re-judge a high-severity-but-unconfirmed candidate.
|
|
758
693
|
|
|
694
|
+
**Enforced:** `validators/validate-run.py` `_validate_final_verification_consistency` fails a report whose `acceptanceBlockers` and `finalVerdict` disagree, so a promoted blocker cannot sit beside an `accepted` verdict and an `accepted` report cannot carry one.
|
|
695
|
+
|
|
759
696
|
### Verdict impact
|
|
760
697
|
|
|
761
698
|
Promoted blockers enter `## 5.8 Acceptance Blockers`; since `accepted` requires zero blockers, the verdict moves to `conditional-accept` / `blocked` automatically. The existing verdict↔blocker consistency validator (`validators/validate-run.py` `_validate_final_verification_consistency`) enforces this unchanged — no new enum or validator.
|
|
@@ -55,6 +55,8 @@ Read-side inspection (`/okstra-inspect`) and scheduling (`/okstra-schedule-gen`)
|
|
|
55
55
|
- If the brief is incomplete, continue with explicit uncertainty markers rather than fabricating confidence.
|
|
56
56
|
- Required roles must not be replaced by unnamed generic parallel workers. Before the final verdict, every selected worker must have either a saved result file or an explicit terminal status with reason. Any attempted worker with status `completed`, `timeout`, or `error` must also have a saved worker prompt history file at its assigned run-level prompt path.
|
|
57
57
|
|
|
58
|
+
**Enforced:** `validators/validate-run.py` `validate_team_state` fails a run whose worker carries a terminal status with no result file or no saved prompt history. The lead's only sanctioned change to the writer's narrative is a mechanical correction ledger — `scripts/okstra_ctl/report_corrections.py` `check_corrections` sets `mechanical` false for any `rewrite` entry, and `okstra agent-prompt apply-corrections` refuses to apply such a ledger.
|
|
59
|
+
|
|
58
60
|
## Worker instruction quality gate (BLOCKING)
|
|
59
61
|
|
|
60
62
|
This gate applies before every initial dispatch, retry, re-verification, critic, and correction, regardless of worker role or transport. The lead performs it on the complete proposed instruction before persisting or sending that instruction.
|
|
@@ -98,10 +100,14 @@ User-utterance interpretation rule:
|
|
|
98
100
|
- If the current phase's outputs are already complete and the user clearly wants to advance, reply with the phase-transition checklist above and the exact next-run command. Wait for explicit user confirmation before any action that belongs to the next phase.
|
|
99
101
|
- If the Next-Phase Pointer target (`nextRecommendedPhase.phase`) is `implementation-planning`, the next run produces a **plan**, not code. The next run after that is `implementation`.
|
|
100
102
|
|
|
103
|
+
**Enforced:** the Forbidden actions column is declared in `prompts/profiles/forbidden-actions.json` and scanned post-hoc by `validators/forbidden_actions.py` `scan_forbidden_actions`, which matches this run's task type through `forbidden_patterns_for` against the commands the run actually executed.
|
|
104
|
+
|
|
101
105
|
## Progress reporting (BLOCKING)
|
|
102
106
|
|
|
103
107
|
A single okstra run frequently spans 30–120 minutes with multi-minute silent windows while workers run; without progress signals the user cannot distinguish "still working" from "hung". Lead MUST emit a single short progress line at each checkpoint below — plain user-facing text in a separate brief message (not buried inside a tool call), one line per checkpoint, format: `PROGRESS: <phase-id> <verb-phrase>`. Emit the line raw — the literal `PROGRESS:` token must begin the line. Do NOT wrap it in inline-code backticks (`` `PROGRESS: ...` ``) or a ```` ``` ```` code fence; markdown wrapping is what the post-hoc conformance validator scrapes around, and raw emit keeps the signal unambiguous.
|
|
104
108
|
|
|
109
|
+
Record each checkpoint with `okstra lead-progress append --project-root <dir> --run-manifest <path> --phase <phase-id>`, then emit the `progressLine` it prints as the user-facing line. The command resolves `leadEventsPath` from the run manifest and writes the checkpoint there — on a host whose adapter declares `sessionAccounting: artifact-only` that ledger is the only place the post-hoc validator can read it, so a conversation line alone leaves no trace and the run is reported as missing the checkpoint. Pass `--worker <role>` on the per-worker checkpoints (it rewrites a phase-specific functional label into the roster role team-state records), `--field NAME=VALUE` for the remaining `key=value` tokens in the order the line carries them, and `--detail <text>` where the line ends in a verb phrase; the wording listed below is the default for the fixed-prose checkpoints.
|
|
110
|
+
|
|
105
111
|
For an `implementation-planning` run whose run manifest declares `activityContractVersion: 1`, record every required activity boundary with `okstra agent-activity append` against the manifest-provided `leadEventsPath`. Model-facing calls pass prose through `--summary-file <md>` and a command through `--command`, `--command-cwd`, `--command-exit-code`, and `--command-output-file <md>`; do not construct `--command-record` JSON. The ordering is fixed: the structured append succeeds first, the matching `PROGRESS:` line is emitted second, and the immediately following `ACTIVITY:` line projects the same structured fields into the conversation language. Do not reconstruct structured activity from conversation text. If the append fails, do not present that activity boundary as completed.
|
|
106
112
|
|
|
107
113
|
The live projection follows this shape:
|
|
@@ -138,7 +144,7 @@ Required checkpoints:
|
|
|
138
144
|
|
|
139
145
|
Do NOT replace them with prose ("Now I'm starting Phase 2..."), do NOT skip a checkpoint because "the previous message already said that", and do NOT batch multiple checkpoints into one. Each line stands alone so the user (or any operator scraping stdout) can timestamp it externally.
|
|
140
146
|
|
|
141
|
-
`okstra-run` surfaces these lines to the user directly;
|
|
147
|
+
`okstra-run` surfaces these lines to the user directly; `okstra lead-progress append` persists them in the selected adapter's declared conformance evidence/event source for post-hoc retrieval.
|
|
142
148
|
|
|
143
149
|
**Enforcement:** the Phase 7 validator (`validators/validate-run.py` → `validate_session_conformance.py`) reads the selected adapter's declared conformance evidence/event source within the run window and fails the run as `contract-violated` when a required checkpoint is missing — including the per-worker `phase-4-dispatch` / `phase-5-collect` lines (which must name each worker's role) and the `phase-batch-cleanup` lines that MUST precede the first `phase-5.5-convergence` round and the `phase-6-synthesis` report-writer dispatch. When the plan-body state file records two or more rounds, `_check_plan_verify_cleanup_checkpoints` additionally requires a `phase-5.5.9-plan-verify` line per round and a `phase-batch-cleanup` between consecutive rounds. For `implementation`, `_check_progress_checkpoints` additionally requires the `phase-5-stage` announcement once any worker is dispatched, and the `phase-5-stage-complete` line once an implementer worker completed. For activity-contract-v1 planning, `_check_activity_contract` validates the structured worker pairs, verification and self-fix counts, user-decision references, and `A-NNN` ordering. `phase-7-teardown` and `complete` fire after validation and are not checked.
|
|
144
150
|
|
|
@@ -166,7 +172,7 @@ The sequence is fixed:
|
|
|
166
172
|
|
|
167
173
|
1. Investigate before asking. Read every cited plan item, worker finding, and `path:line` the question depends on. You are not ready to ask while any option's outcome cannot be named as a concrete change: the extra work it creates, the already-decided thing it reverses, and which files or stages it touches. If a citation cannot be read, say so in the question; do not invent the missing fact.
|
|
168
174
|
2. Emit `PROGRESS: user-confirm <C-NNN> <the question, one line>` with the id the row would carry.
|
|
169
|
-
3. Ask in the user's language through the selected adapter's `prompt_user` mapping. Do not lead with `C-NNN`, `Kind`, `Blocks`, or `expectedForm`. The question body is why this is being asked, what is already decided, the fork, and each option as "if you pick this, then …". Keep the report-owned impact axes (reach or scope, added work, direction change). One question at a time. When that mapping names a native question function and the option count fits the relay `nativeLimits`, call that function with `{label, description}` options. Do not print a numbered list in chat while the native tool is available. Numbered text is only for a text-only relay or a prompt that does not fit `nativeLimits`.
|
|
175
|
+
3. Ask in the user's language through the selected adapter's `prompt_user` mapping. Do not lead with `C-NNN`, `Kind`, `Blocks`, or `expectedForm`. The question body is why this is being asked, what is already decided, the fork, and each option as "if you pick this, then …". Keep the report-owned impact axes (reach or scope, added work, direction change). One question at a time. When that mapping names a native question function and the option count fits the relay `nativeLimits`, call that function with `{label, description}` options. Do not print a numbered list in chat while the native tool is available. Numbered text is only for a text-only relay or a prompt that does not fit `nativeLimits`. Emit the question as the last thing in that turn — no restatement, rationale, or status line after the call. Anything emitted after it renders below the question and separates it from the user's answer.
|
|
170
176
|
4. On an answer — record the raw text in the row's `userInput`, set `status: answered` and `userConfirmation: asked-and-answered`, and apply the selected disposition in this run.
|
|
171
177
|
5. Only when asking fails does the row stay open: `asked-awaiting` when the user has not answered, `deferred-no-interactive-session` when this run has no user to ask.
|
|
172
178
|
|
|
@@ -221,6 +227,8 @@ Lead MUST dispatch Edit/Write-bearing work only through that executor binding: u
|
|
|
221
227
|
|
|
222
228
|
Executor is chosen at run-prep time via `--executor <claude|codex|antigravity>` (or `OKSTRA_DEFAULT_EXECUTOR`, fallback `claude`); the model used by the executor is taken from the corresponding worker model flag (`--claude-model` / `--codex-model` / `--antigravity-model`). For CLI-backed executors, the underlying file mutation happens inside the executor CLI's own auto-edit mode (e.g. `codex exec --sandbox danger-full-access`), not through the lead runtime's `write_artifact` operation.
|
|
223
229
|
|
|
230
|
+
**Enforced:** every non-executor assignment is dispatched under `writePolicy.sourcePolicy.mode` `source-readonly`, and `scripts/okstra_ctl/execution_mutation_audit.py` `_source_changes` compares the before/after snapshots and reports a read-only worker that mutated source; `assert_compatible_batch` refuses a batch that mixes a mutating worker with read-only ones.
|
|
231
|
+
|
|
224
232
|
#### Task worktree (BLOCKING for every task-type)
|
|
225
233
|
|
|
226
234
|
Okstra provisions dedicated `git worktree`s at run-prep time. Lead, the Executor, and every verifier MUST treat the provisioned worktree as the canonical working directory regardless of task-type.
|
|
@@ -235,6 +243,7 @@ Okstra provisions dedicated `git worktree`s at run-prep time. Lead, the Executor
|
|
|
235
243
|
- `project_root` is already inside a non-main worktree (the run reuses the caller's worktree to avoid nesting).
|
|
236
244
|
- `project_root` is not inside a git repository at all.
|
|
237
245
|
- **Failure mode**: any other error during `git worktree add` (on-disk path collision not tracked by the registry, branch collision with a different task-key, detached-HEAD with no SHA) raises `PrepareError`. Re-run after manually removing the stale path/branch — the worktree is intentionally not garbage-collected.
|
|
246
|
+
- **Enforced:** `scripts/okstra_ctl/dispatch_state.py` `worktree_path` resolves the provisioned worktree for every task-type and materialization writes it into each worker prompt as the `**Worktree:**` anchor; `scripts/okstra_ctl/worker_prompt_contract.py` `_REQUIRED_TARGET_PREFIXES` requires that anchor on the prompts that carry it, and `scripts/okstra_ctl/execution_mutation_audit.py` `_source_changes` audits mutations against that root.
|
|
238
247
|
- **Lifecycle**: kept after the run for follow-up phases, manual PR authoring, rollback verification, and `final-verification`. Manual cleanup: `git -C <main-worktree> worktree remove <path>` then `git -C <main-worktree> branch -D <branch>`; remove the corresponding (task- or stage-) key from the registry by hand.
|
|
239
248
|
|
|
240
249
|
## Phase 1: Task-bundle intake and required reading order
|
|
@@ -312,6 +321,8 @@ For `improvement-discovery`, Lead records `## Primary Pass Assignments` in the P
|
|
|
312
321
|
4. Persist the setup outcome in team-state using the existing fields required by that backend.
|
|
313
322
|
5. Emit the canonical `PROGRESS: phase-3-team-create <adapter-specific-status>` checkpoint. The phase id remains stable for artifact compatibility; only the adapter-owned verb phrase varies.
|
|
314
323
|
|
|
324
|
+
**Enforced:** `validators/validate_session_conformance.py` `_check_progress_checkpoints` requires the `phase-3-team-create` checkpoint once any worker was dispatched, and `validate-run.py` `validate_team_state` reads the setup outcome this phase persists.
|
|
325
|
+
|
|
315
326
|
### Phase 4 / Phase 5 — Dispatch, await, and error-log recording
|
|
316
327
|
|
|
317
328
|
For each selected worker assignment, persist the exact prompt history, emit the per-worker Phase 4 checkpoint, and call `dispatch_worker(assignment, promptPath)` through the selected adapter — the dispatch hands the worker the persisted prompt's path, never a re-inlined copy of its body. Then call `await_workers(handles)`. A dispatch acknowledgement or process/pane creation is never completion: verify the terminal status, Result Path, and worker-results audit path required by `team-contract` before emitting the Phase 5 collection checkpoint.
|
|
@@ -343,7 +354,7 @@ For deterministic Codex/Antigravity provider processes: if the CLI returns non-z
|
|
|
343
354
|
```bash
|
|
344
355
|
okstra error-log append-observed \
|
|
345
356
|
--out <absolute-errors-log-path-from-launch-prompt> \
|
|
346
|
-
--task-key <taskKey> --phase
|
|
357
|
+
--task-key <taskKey> --phase <task-type> \
|
|
347
358
|
--agent codex-worker --agent-role worker --model <model> \
|
|
348
359
|
--error-type cli-failure \
|
|
349
360
|
--command "okstra-codex-exec.sh <project-root> <model> <prompt-path>" \
|
|
@@ -387,6 +398,8 @@ If convergence is disabled, `seed`/`finalize` produce the auto-disabled final st
|
|
|
387
398
|
|
|
388
399
|
**REQUIRED RESOURCE:** Read [report-writer](./report-writer.md) for report ownership, dispatch, assembly, and Phase 7 rules.
|
|
389
400
|
|
|
401
|
+
**Enforced:** `validators/validate_session_conformance.py` `_check_progress_checkpoints` requires the `phase-6-synthesis` checkpoint once the report writer was dispatched, `validators/validate-run.py` `validate_team_state` requires that dispatch's result and prompt history, and `scripts/okstra_ctl/report_corrections.py` `check_corrections` keeps the narrative the writer's — the obligations this resource carries. Loading the document itself is not separately checked.
|
|
402
|
+
|
|
390
403
|
### Authoring ownership (BLOCKING)
|
|
391
404
|
|
|
392
405
|
If `Report writer worker` is in the selected roster (`recommendedWorkers` / `resultContract.requiredWorkerRoles`), Lead dispatches it to author only `report-writer-narrative-<task-type>-<seq>.md`, its worker-result pointer, and its audit sidecar. The worker may read the complete run context but cannot write `final-report-*.data.json` or another role's ledger. After every required input exists, Phase 7 runs report assembly, which validates the role-owned inputs and atomically publishes the report record once. Phase 7 then renders the human HTML from that record. Contract v2 artifacts remain readable but no new run writes them. **Enforced:** report-writer dispatch completion paths in `scripts/okstra_ctl/dispatch_state.py`, granted artifacts in `scripts/okstra_ctl/dispatch_core.py`, and `scripts/okstra_ctl/report_assembly.py` `assemble_report`.
|
|
@@ -451,8 +464,8 @@ The detailed persistence sequence lives in [report-writer](./report-writer.md).
|
|
|
451
464
|
|
|
452
465
|
Order of operations:
|
|
453
466
|
|
|
454
|
-
1. Run `okstra report-finalize ...`. Contract v3 collects usage, assembles the role-owned inputs into `data.json` once, checks the source, renders views, persists follow-ups, validates the run, and performs eligible teardown in order.
|
|
455
|
-
2. When `meta.reportLanguage` is not `en`, first run only `token-usage`, `project-activity`, and `check-source`. Dispatch the translator against that assembled record, then resume with only `render-views`, `spawn-followups`, `validate-run`, and `teardown-stages` so assembly is not repeated.
|
|
467
|
+
1. Run `okstra report-finalize ...`. Contract v3 collects usage, assembles the role-owned inputs into `data.json` once, checks the source, renders views, persists follow-ups, validates the run, records this run's conclusion into the task-group's `group-context.md` for the group's sibling tasks, and performs eligible teardown in order. When the task's pointer is terminal and the group has a task not yet started, the result's `nextInGroup` names it and `nextCommand` closes on starting it from its brief — the group's start order is the brief ordinal; the briefs' Related Task Graph edges are quoted as `waits for` and do not reorder it.
|
|
468
|
+
2. When `meta.reportLanguage` is not `en`, first run only `token-usage`, `project-activity`, and `check-source`. Dispatch the translator against that assembled record, then resume with only `render-views`, `spawn-followups`, `validate-run`, `record-group-memory`, and `teardown-stages` so assembly is not repeated.
|
|
456
469
|
|
|
457
470
|
Keep the assigned worker prompt history paths stable in `team-state`, `run-manifest`, and `task-manifest`. Do not rewrite prompt artifacts to `/tmp` or omit prompt metadata for attempted workers.
|
|
458
471
|
|
|
@@ -470,6 +483,7 @@ For every other task type:
|
|
|
470
483
|
|
|
471
484
|
- Pointer `status: ready` → `/okstra-run` for that `phase`. Quote the pointer's `rationale` — that sentence is this run's report saying why that phase comes next, and it is the analysis the user asked for.
|
|
472
485
|
- Pointer `status: terminal` → the lifecycle ends here. Say the task is finished and quote the pointer's `rationale`. Do not propose a run and do not send the user to `/okstra-inspect`: the decision is already made, so there is nothing to inspect. If this run registered follow-up tasks, name them and the command that starts one.
|
|
486
|
+
- Pointer `status: terminal` and the result carries `nextInGroup` → the task is finished and the task-group has a task not yet started: say this task is finished, then close on `/okstra-run` for `nextInGroup.briefId` from its `brief` path — the group's start order is the brief ordinal, and `nextCommand.note` already names the task and its brief.
|
|
473
487
|
- Pointer `status: blocked` → the pointer's `rationale` names what is in the way and what to run. Quote it and issue that command — `/okstra-user-response` for the `C-NNN` ids it lists, or `/okstra-run` for the phase it names.
|
|
474
488
|
- Pointer `status: pending` with a `rationale` → the comparison finished and the user's own choice is what comes next. Quote the `rationale` and issue the command it names. Do not re-run the phase that just completed, and do not send the user to `/okstra-inspect`: a finished phase has nothing to inspect, and re-running it discards the result the user is being asked to choose from.
|
|
475
489
|
- Phase 7 `validate-run` failed → one line naming the blocking cause, then `/okstra-run` to re-run this phase with the recorded sidecar. Only a failure matching the blocking allowlist (`okstra_ctl.blocking_checks`) reaches this row; every other finding was demoted to an advisory, printed as `validate-run: advisory — <finding>`, and the run passed. Name those advisories in one line and take the command from the matching pointer row above — an advisory is not a reason to re-run.
|
|
@@ -477,7 +491,11 @@ For every other task type:
|
|
|
477
491
|
|
|
478
492
|
When the host native picker is available and two of those rows could apply, ask with that picker (recommended first). Do not end the turn after the status dump.
|
|
479
493
|
|
|
480
|
-
**
|
|
494
|
+
**Cite run-artifact paths, do not assemble them.** The `report-finalize` result's `reportPaths` carries this run's `humanReport`, `reportRecord`, `teamState`, and `renderFullCopy` command, each already rooted at the project (`.okstra/tasks/<task-group>/<task-id>/runs/...`). Every other run-artifact path the reply cites — the resume command among them — comes from the launch prompt's `## Manifests` / `## Run Paths` lists, which are rooted the same way. A path you compose from a `runs/<task-type>/...` pattern instead is identical across every task of that task-type, so it names no task and does not resolve from the project root either.
|
|
495
|
+
|
|
496
|
+
**Write every file you name to the user as a markdown link** — `[<what it is>](<path>)`, the path inside the parentheses. The host renders the reply as markdown, so that is the only form the user can click; a path in backticks or bare prose is text they have to copy out and open by hand. `reportPaths.markdown` already carries this run's report, report record, and team state in that form — paste those strings. Apply the same shape to every other file the reply names (a worker result, the errors log, a brief), using the project-rooted path from the launch prompt's `## Manifests` / `## Run Paths`. Commands stay in backticks; only files become links.
|
|
497
|
+
|
|
498
|
+
**Enforced:** `scripts/okstra_ctl/report_finalize.py` `_closeout_report_paths` builds the task-qualified paths and `closeout_command` builds the close, so neither is re-derived; `validators/validate_session_conformance.py` `_check_progress_checkpoints` requires the `phase-7-persist` checkpoint this phase opens with.
|
|
481
499
|
|
|
482
500
|
The run-level error log lives at `<runDir>/logs/errors-<task-type>-<seq>.jsonl`. It is informational. Its presence or absence does not affect the final verdict. Do not block report writing on it.
|
|
483
501
|
|
|
@@ -25,6 +25,8 @@ Phase 4 workers produce independent analyses (Findings F-001…)
|
|
|
25
25
|
|
|
26
26
|
Plan-body verification MUST NOT replace, precede, or be conflated with the Phase 5.5 finding convergence above. They are two distinct rounds with different inputs (findings vs. consolidated plan body), different ID schemes (`F-*` vs. `P-*`), and different state files.
|
|
27
27
|
|
|
28
|
+
**Enforced:** `validators/validate-run.py` `_validate_queue_mutual_exclusion` fails a run whose finding queue carries a `P-*` id or whose plan-item queue carries an `F-*` / `C-*` id — the two rounds cannot be merged without tripping it.
|
|
29
|
+
|
|
28
30
|
## MUTUAL EXCLUSION (BLOCKING)
|
|
29
31
|
|
|
30
32
|
The finding queue (Phase 5.5, [convergence](./convergence.md)) and the plan-item queue (this contract) are **disjoint**:
|
|
@@ -312,6 +314,8 @@ which is what sends the item into a self-fix round that can only paper over it.
|
|
|
312
314
|
|
|
313
315
|
Plan-body verification only supports **lightweight mode** (defined in [convergence](./convergence.md) §"Verification Mode"). `full-reanalysis` is not meaningful here because the "original source materials" for a plan item are the worker's own analysis plus the lead-mediated synthesis — there is no independent ground truth to re-read. The manifest's top-level `verificationMode` is ignored for this round; lightweight is always used.
|
|
314
316
|
|
|
317
|
+
**Enforced:** `scripts/okstra_ctl/convergence_engine.py` `_parse_config` rejects a `verificationMode` outside `_VERIFICATION_MODES`, and the plan-body round is seeded with `lightweight` rather than the manifest value.
|
|
318
|
+
|
|
315
319
|
Exception for `P-Req-*`: verifiers still MUST NOT re-open the original task brief for this round, but they MUST compare the requirement text embedded in the `Requirement Coverage` row with the cited Option / Stage / Step in the draft plan. A row is not sound merely because it says `covered`; the cited plan item must actually satisfy the row's stated requirement. A `documented-deviation` never earns automatic `AGREE`: verify all four parts independently — the original requirement, the concrete alternative in `coveredBy`, every `decisionRefs` target, and the `approvalDisposition`. `accepted` requires a referenced user-confirmed clarification; `blocked C-NNN` requires that same-report clarification to be open and block approval.
|
|
316
320
|
|
|
317
321
|
## Adversarial plan-body posture
|
|
@@ -385,7 +389,7 @@ round before any host or provider process starts.
|
|
|
385
389
|
**Score the gate with `okstra plan-verify`, never by hand (BLOCKING).** Once this round's verdicts are in the data.json, lead runs
|
|
386
390
|
|
|
387
391
|
```
|
|
388
|
-
okstra plan-verify --report <
|
|
392
|
+
okstra plan-verify --report <the launch prompt's `## Run Paths` → `Final report` value>
|
|
389
393
|
```
|
|
390
394
|
|
|
391
395
|
and records what it returns: `gate.recomputed` is the round's gate value, `gate.blockedBy` its `gateBlockedBy` causes, `gate.blockingItems` the items that block. The same call runs every plan-body check a round can be judged on alone — provenance, fixability, subject substance, self-fix grouping, round recording, clarification matching, state-file rounds — and exits 2 with `failures[]` when the round is not contract-clean. **A round is not complete while that exit code is non-zero.**
|
|
@@ -670,12 +674,15 @@ What concrete false-positive input, failure ordering, or omitted dependency
|
|
|
670
674
|
would make this plan item incorrect even if its happy path succeeds?
|
|
671
675
|
```
|
|
672
676
|
|
|
673
|
-
Even for `AGREE`, the worker
|
|
674
|
-
the supplied `payload` excludes it in the verdict note
|
|
677
|
+
Even for `AGREE`, the worker records the considered counterexample and why
|
|
678
|
+
the supplied `payload` excludes it in the verdict note — that is what the
|
|
679
|
+
explanation below has to contain. If capability,
|
|
675
680
|
credential, network, or service state prevents verification, record
|
|
676
681
|
`UNVERIFIABLE` for that item. It is persisted as `verification-error`, excluded from both numerator and denominator, and must not be converted to `DISAGREE`.
|
|
677
682
|
Every verdict requires a non-empty explanation.
|
|
678
683
|
|
|
684
|
+
**Enforced:** `scripts/okstra_ctl/convergence_engine.py` `_parse_vote` rejects a vote whose `explanation` is empty and whose verdict is outside `_VERDICTS`; `classify_collaborative_round` drops `verification-error` votes from both numerator and denominator.
|
|
685
|
+
|
|
679
686
|
For `P-Req-*` items, compare only the requirement text embedded in the row
|
|
680
687
|
against the cited plan item(s). Do not open the original brief, but do reject
|
|
681
688
|
coverage rows that cite no concrete option/stage/step or cite a plan item that
|
|
@@ -24,12 +24,19 @@ Report assembly reads the role-owned inputs, validates them, derives links and s
|
|
|
24
24
|
|
|
25
25
|
An active clarification exists only in `activeClarifications[]`. A decision carried from a previous run exists only in `carriedDecisions[]`; do not recreate it as an active question.
|
|
26
26
|
|
|
27
|
-
Prepare seeds `carriedDecisions[]` when it creates the ledger: every clarification the run's carry-in record answered or resolved, plus every row that record's user-responses sidecars answered, arrives carried (`scripts/okstra_ctl/approval_decisions.py` `seed_carried_decisions`). A carried plan row's `requirementCoverage[].decisionRefs` may still name a `C-NNN` the carry-in record does not answer — a decision from an older run. Carry it before assembly: `okstra approval-decision carry --ledger <approvalDecisionsPath> --from-responses <instruction-set/clarification-response.md> --clarification-id C-NNN` — repeat `--clarification-id` to take several in one call. That bundle is the source of truth for an earlier run's answer: it is task-level and cumulative, so no prior run seq has to be located, and each response section names the report that posed the question, which is where the row's `statement`, `expectedForm`, and options come from. A carried row lands as `answered`, not `resolved` — it was resolved in another run, and `resolution.checkRefs` names *this* run's activity rows. Carrying an answer also obliges a `supersessionLedger` entry for it.
|
|
27
|
+
Prepare seeds `carriedDecisions[]` when it creates the ledger: every clarification the run's carry-in record answered or resolved, plus every row that record's user-responses sidecars answered, arrives carried (`scripts/okstra_ctl/approval_decisions.py` `seed_carried_decisions`). The carry-in record is the `--clarification-response` file; for a new plan it is the option-selection record `--selected-direction` names, and for an implementation run the approved plan `--approved-plan` names (`render._carry_in_source`) — the same pointer assembly writes to `clarificationCarryIn.sourceFile`, so the page links those ids to the prior run's page. A carried plan row's `requirementCoverage[].decisionRefs` may still name a `C-NNN` the carry-in record does not answer — a decision from an older run. Carry it before assembly: `okstra approval-decision carry --ledger <approvalDecisionsPath> --from-responses <instruction-set/clarification-response.md> --clarification-id C-NNN` — repeat `--clarification-id` to take several in one call. That bundle is the source of truth for an earlier run's answer: it is task-level and cumulative, so no prior run seq has to be located, and each response section names the report that posed the question, which is where the row's `statement`, `expectedForm`, and options come from. A carried row lands as `answered`, not `resolved` — it was resolved in another run, and `resolution.checkRefs` names *this* run's activity rows. Carrying an answer also obliges a `supersessionLedger` entry for it.
|
|
28
28
|
|
|
29
29
|
Each decision option has `role`, `answer`, `rationale`, `disposition`, `reach`, optional `scopeEffects`, `addedWork`, and `directionChange`. `reach` is exactly one of `in-repo` or `cross-repo`. `scopeEffects` may contain `new-schema` and `deferrable`. A `correctness-critical` option cannot use `select` or `accept-risk`; a `noncritical-dissent` option cannot use `select`.
|
|
30
30
|
|
|
31
31
|
Resolution `checkRefs` name existing `A-NNN` activity rows. Those activity rows carry `clarificationRefs[]`; their `planItemIds[]` let report assembly derive the reverse plan-item links. Do not store copied plan or activity identifiers in `approvalContext`.
|
|
32
32
|
|
|
33
|
+
## Field content rules
|
|
34
|
+
|
|
35
|
+
Two rules the assembled record is checked against, both learned from the 2026-09-05 audit of shipped pages:
|
|
36
|
+
|
|
37
|
+
- **An empty value is empty, never a null literal.** A prose field with nothing to say is omitted or left `""`. `"None"`, `"null"`, `"undefined"` and `"NaN"` are serialisation artefacts the page prints as text (`tradeoffMatrix` cells in nlpvibe planning-009…014). The lowercase `none` an option's `addedWork`/`directionChange` uses to mean "nothing" is a contract token and stays. **Enforced:** `validators/validate-run.py` `_validate_no_null_literals_in_prose` fails a prose field holding one of those literals.
|
|
38
|
+
- **Citations name rows the reader can reach.** `evidenceRefs`, `supportingEvidence` and the other citation lists carry the record's own ids (`E-`, `CV-`, `D-`, `EA-`, `C-`, `A-`) or a `path:line`. A worker's own finding number (`F-NNN`) reaches the reader only through a promoted `evidence.primary[]` row whose `sourceItems` name it as `<worker>:F-NNN` — the page links a bare `F-NNN` to that row when exactly one row carries it, and leaves it as dead text otherwise. **Enforced:** `_unbridged_worker_finding_refs` in `validators/validate-run.py` warns on every bare `F-NNN` no `sourceItems` entry carries; the narrative rule `_validate_no_opaque_id_references` already fails one in a reader-facing sentence.
|
|
39
|
+
|
|
33
40
|
## Report-writer dispatch
|
|
34
41
|
|
|
35
42
|
For report contract 3.0, prompt materialization first freezes one report synthesis
|
|
@@ -64,7 +71,9 @@ Materialize the duty prompt with `okstra agent-prompt materialize --audience rep
|
|
|
64
71
|
|
|
65
72
|
The errors sidecar anchor reserves the runtime-owned write-artifact path used by dispatch validation. No model-authored error JSON file is part of report-writer dispatch; failures use the typed error-log command from the worker error contract.
|
|
66
73
|
|
|
67
|
-
|
|
74
|
+
After materialization, generate a v2 jobs file with `okstra agent-prompt jobs --project-root <root> --run-manifest <manifest> --dispatch-kind report-writer --metadata <returned-meta.json> --out <run-state>/report-writer-jobs.json`. Pass the generated file through the selected deterministic dispatcher's `--jobs-file`. The generator reads the narrative and worker-result headers and validates their existing manifest contract. For host-native dispatch, use `okstra agent-prompt record-dispatch` and `link-result`; deterministic dispatchers record their own invocation lifecycle.
|
|
75
|
+
|
|
76
|
+
When `terminalBackend` is `cmux-pane`, use `okstra team dispatch --project-root <root> --run-manifest <manifest> --jobs-file <jobs-file>`. For a CLI wrapper, use `okstra worker-dispatch --project-root <root> --run-manifest <manifest> --jobs-file <jobs-file>`. Both paths use `okstra team await` to collect completion.
|
|
68
77
|
|
|
69
78
|
The pointer record contains the narrative and audit paths. Completion never depends on the final record because assembly runs after writer completion.
|
|
70
79
|
|
|
@@ -114,7 +123,7 @@ For historical schema-v1 Markdown only, the following heading table remains a re
|
|
|
114
123
|
|
|
115
124
|
**Enforced:** `okstra_ctl.report_finalize.V3_STEP_ORDER` is the order — `report-finalize` runs the steps from that tuple, so the sequence cannot be reordered by a caller. Running the steps by hand is what this rule forbids, and that path is not reachable through the CLI.
|
|
116
125
|
|
|
117
|
-
Do not run the
|
|
126
|
+
Do not run the eight steps below manually. Invoke `okstra report-finalize`; contract 3.0 runs them in this order:
|
|
118
127
|
|
|
119
128
|
1. **`token-usage`** — collect usage into team state without touching the final record.
|
|
120
129
|
2. **`project-activity`** — report assembly validates every owner input and publishes the final record once.
|
|
@@ -122,13 +131,14 @@ Do not run the seven steps below manually. Invoke `okstra report-finalize`; cont
|
|
|
122
131
|
4. **`render-views`** — render the Markdown reading copy and human HTML.
|
|
123
132
|
5. **`spawn-followups`** — materialize registered follow-up tasks.
|
|
124
133
|
6. **`validate-run`** — validate the record, views, run manifest, and team state.
|
|
125
|
-
7. **`
|
|
134
|
+
7. **`record-group-memory`** — write this run's conclusion (headline, decisions, watch-outs, open follow-ups, record path, next phase) and the group's start order into the task-group's `group-context.md` okstra region, creating the file when the group has none; skipped, like teardown, when an earlier step failed. Sibling tasks read it as `## Task-Group Memory`.
|
|
135
|
+
8. **`teardown-stages`** — remove eligible stage worktrees after successful validation.
|
|
126
136
|
|
|
127
137
|
After `report-finalize` returns, the lead — not the report writer — closes the run with the launch prompt's User closeout: one command the user can run now.
|
|
128
138
|
|
|
129
139
|
### Before `report-finalize`: the translation sidecar
|
|
130
140
|
|
|
131
|
-
Never dispatch the translator before report assembly and `check-source`. For a non-English human report, first run `report-finalize --only token-usage --only project-activity --only check-source`; the extraction command refuses to build a work list from a non-English source. Then dispatch the translator worker with `okstra agent-prompt materialize --audience translator`, `okstra agent-prompt record-dispatch`, `okstra worker-dispatch --
|
|
141
|
+
Never dispatch the translator before report assembly and `check-source`. For a non-English human report, first run `report-finalize --only token-usage --only project-activity --only check-source --only render-views`; the extraction command refuses to build a work list from a non-English source. `render-views` belongs in that first call: the HTML view renders from the report body and never reads the translation sidecar, so rendering it early costs one rerun and means a translator that never starts does not also cost the run its only human-readable artifact. Then dispatch the translator worker with `okstra agent-prompt materialize --audience translator --worker-id translator --dispatch-kind translator --assignment-ref translator --result <run>/worker-results/translator-translations-<task-type>-<seq>.md --audit-source <run>/worker-results/translator-worker-<task-type>-<seq>.md`, `okstra agent-prompt record-dispatch`, `okstra worker-dispatch --workers translator`, and `okstra agent-prompt link-result`. `--result` is the translator's own report and `--audit-source` the worker result the audit sidecar derives from; the two must differ, and the materializer refuses a translator prompt without a distinct `--audit-source` because `worker-dispatch` reads both from the prompt's `**Result Path:**` and `**Worker Result Path:**` anchors and refuses a prompt missing the second. A translator prompt that failed a pre-dispatch gate is re-materialized under the **same** invocation id with `--replace-undispatched`; a second invocation id leaves two undispatched reservations for one run, there is no command that withdraws a reservation, and `worker-dispatch` refuses to choose between them. **Enforced:** `_materialize_run` in `scripts/okstra_ctl/agent/prompt_cli/materialize.py`; `_translator_job_from_reservation` in `scripts/okstra_ctl/dispatch_core.py`. Resume with `report-finalize --only render-views --only spawn-followups --only validate-run --only record-group-memory --only teardown-stages`; do not assemble the record a second time. `render-views` is idempotent and overwrites the view it already wrote, this time with the translation overlaid. When the translator cannot run at all — a host approval gate, an unavailable provider — say so and finish the sequence anyway: the run keeps an English view rather than no view, and `report-finalize` inserts the missing `render-views` itself when a `--only validate-run` call would otherwise validate a report that has none (**Enforced:** `okstra_ctl.report_finalize._with_view_the_validator_reads`).
|
|
132
142
|
|
|
133
143
|
## Routing pointer
|
|
134
144
|
|
|
@@ -68,11 +68,11 @@ Only workers selected from `recommendedWorkers` in `task-manifest.json` and `res
|
|
|
68
68
|
|
|
69
69
|
## Worker Prompt Composition
|
|
70
70
|
|
|
71
|
-
`okstra_ctl.initial_prompt_materialization` is the canonical owner of roster-derived initial prompt rendering, validation, and immutable publication. Code-backed `dispatch_worker` mappings leave missing roster prompts to `materialize_initial_prompts()` and pass the selected adapter's declared `initialPromptDeliveryMode`; they do not compose or overwrite those prompts themselves. Native in-process dispatch uses the same headers and the
|
|
71
|
+
`okstra_ctl.initial_prompt_materialization` is the canonical owner of roster-derived initial prompt rendering, validation, and immutable publication. Code-backed `dispatch_worker` mappings leave missing roster prompts to `materialize_initial_prompts()` and pass the selected adapter's declared `initialPromptDeliveryMode`; they do not compose or overwrite those prompts themselves. That declaration reaches the materializer through the run manifest's `leadAdapter.initialPromptDeliveryMode`, which render copies from the host descriptor — or from the dispatch adapter that took worker dispatch over, such as cmux. Every adapter currently declares `eager-include`: `lazy-path-reference` drops the clarification-answer and fix-carry blocks out of an implementation prompt without listing them as resource paths, so it is not a mode an adapter may declare until that gap is closed. **Enforced:** `okstra_ctl.dispatch_core._initial_prompt_delivery_mode` refuses a manifest declaring anything outside `okstra_ctl.initial_prompt_materialization.PromptDeliveryMode`. Native in-process dispatch uses the same headers and the same declared mode, persists the prompt before dispatch, and remains outside reverify and critic handling. The native dispatch payload is a summon message carrying the persisted prompt path and the instruction to read that document in full — never the prompt body itself (the host adapter's `dispatch_worker` row owns the exact summon wording).
|
|
72
72
|
|
|
73
73
|
Every initial worker instruction and result-missing retry also passes [okstra-lead-contract](./okstra-lead-contract.md) "Worker instruction quality gate" before it is persisted or sent. Prompt-shape validation does not replace that content check: the lead verifies current paths, exact constrained values, role boundaries, and completion conditions against their authoritative artifacts.
|
|
74
74
|
|
|
75
|
-
Every worker prompt MUST start with the anchor headers rendered by `okstra_ctl.worker_prompt_headers.worker_prompt_headers()` (the generating SSOT — never hand-author or reorder them). Every persisted initial prompt also carries exactly one non-empty `**Prompt Delivery Mode:** <mode>` header whose value is `eager-include` or `lazy-path-reference`. Redispatch reuses that persisted initial prompt byte-for-byte
|
|
75
|
+
Every worker prompt MUST start with the anchor headers rendered by `okstra_ctl.worker_prompt_headers.worker_prompt_headers()` (the generating SSOT — never hand-author or reorder them). Every persisted initial prompt also carries exactly one non-empty `**Prompt Delivery Mode:** <mode>` header whose value is `eager-include` or `lazy-path-reference`. Redispatch reuses that persisted initial prompt byte-for-byte, with one exception it performs for you: a prompt published before an anchor the current contract renders is rewritten in place under the same invocation id, and only while no dispatch row names that invocation. The generated absolute `**Audit sidecar path:**` is derived by `audit_sidecar_rel()` from the canonical worker result path; workers write to that header and never synthesize a `runs/<task-type>/...` destination. Their meaning and extraction rules for workers are owned by `okstra_ctl.worker_prompt_headers.worker_prompt_headers()`. Phase-specific extra headers (final-verification target snapshot, improvement-discovery grilling log) are emitted there too. The `**Worktree:**` anchor is not phase-specific: `okstra_ctl.initial_prompt_materialization._worktree_anchor_lines` renders it into every initial prompt whose run has a provisioned worktree, so an analysis worker anchors its reads to the same tree the executor mutates. **Enforced:** `okstra_ctl.initial_prompt_materialization._anchor_predates_this_contract` republishes a prompt that predates the anchor rather than refusing the dispatch.
|
|
76
76
|
|
|
77
77
|
The `**Coding preflight pack:**` anchor is emitted only for the `implementation-executor` and `implementation-verifier` audiences. An implementation report-writer never receives it. A pane display role named `verifier` during `final-verification` is still an initial analysis worker; it neither receives implementation conventions nor becomes a Phase 5.5 reverify dispatch. Reverify is identified by the `-reverify-r<N>-` prompt/result path contract in [convergence](./convergence.md).
|
|
78
78
|
|
|
@@ -149,9 +149,11 @@ One selector serves every worker kind: `--team-state <path> --worker <id>`, repe
|
|
|
149
149
|
| `livenessMode` | Backend | What it reads |
|
|
150
150
|
|---|---|---|
|
|
151
151
|
| `audit-heartbeat` | in-process dispatch (including `claude-worker` and `report-writer-worker`) | the row's `auditSidecarPath` — its newest `- PROGRESS:` heartbeat, measured from `startedAt` |
|
|
152
|
-
| `wrapper-status` | CLI wrapper (`codex-worker` / `antigravity-worker`) | the row's `promptPath`,
|
|
152
|
+
| `wrapper-status` | CLI wrapper (`codex-worker` / `antigravity-worker`) | canonical wrapper artifacts resolved by `worker-liveness` from the row's `promptPath`, measured from `startedAt` |
|
|
153
153
|
|
|
154
|
-
|
|
154
|
+
Only CLI wrappers write wrapper logs and status sidecars; in-process workers report through their audit sidecar. Each grace begins at the persisted `startedAt`, not prompt-materialization time or an artifact's mtime.
|
|
155
|
+
|
|
156
|
+
For diagnostics, open the dispatch record's `statusSidecarPath` and follow its `log_path` value. `wrapper_status.log_path_for_prompt` removes the prompt's `.md` suffix for the log (`p.log`), while the status path keeps it (`p.md.status.json`). If the status sidecar is absent, use `worker-liveness` or `team await`; do not infer a dead worker from a guessed log filename.
|
|
155
157
|
|
|
156
158
|
This is a transport-adapter liveness choice only. It does not change the reducer's worker identity or its verification responsibility: both remain bound to the registered worker instance and canonical artifacts.
|
|
157
159
|
|
|
@@ -166,21 +168,21 @@ After each worker attempt returns (regardless of role), Lead MUST verify the can
|
|
|
166
168
|
- The deterministic provider process returned an explicit `*_RESULT_MISSING` sentinel.
|
|
167
169
|
- The result file is absent at the resolved absolute path even though the worker returned without a `*_RESULT_MISSING` sentinel — for example, claude-worker returned its final assistant message but never persisted the artifact, or the wrapper exited 0 and the codex/antigravity sub-agent forwarded raw stdout despite the contract.
|
|
168
170
|
- The result file exists but cannot be parsed (frontmatter unreadable, sections 1–5 entirely missing). A truncated file in the middle of section 5 is NOT covered here — it goes to the validator's regular `error` path, not the retry path.
|
|
169
|
-
- `okstra worker-liveness --team-state <path> --worker <id>` reports a **CLI-wrapper** worker
|
|
171
|
+
- `okstra worker-liveness --team-state <path> --worker <id>` reports a **CLI-wrapper** worker `did-not-launch` after its launch grace. The probe resolves both wrapper artifacts through `wrapper_status`; a missing guessed filename is not this verdict. Use the persisted probe outcome before retrying.
|
|
170
172
|
- `okstra worker-liveness --team-state <path> --worker <id>` reports an **in-process** worker `stalled` — its registered audit sidecar's newest `- PROGRESS:` heartbeat is older than the cadence budget, or the sidecar carries no heartbeat at all. This is the in-process equivalent of the CLI wrappers' idle watchdog: the wrapper reaps a silent CLI itself, but nothing reaped a silent in-process worker until its deadline. The same selector serves both worker kinds — the sidecar is reused when a worker is re-dispatched, so the probe needs the row's `startedAt` to tell the previous attempt's last heartbeat apart from this dispatch's silence. A budget breach is not a verdict on its own: the probe re-reads the sidecar after half that stage's budget and reports `stalled` only when the newest heartbeat has not advanced, so an unhealthy verdict costs that confirmation window before it returns. That window is what separates a worker inside one long uninterruptible tool call — which cannot append a heartbeat at all — from a worker that died, and measuring it is the probe's job, not yours: do not second-guess a verdict by checking mtimes yourself. Tune it with `--stall-confirm <seconds>`; `0` restores the immediate verdict.
|
|
171
173
|
- The result file exists but its audit sidecar does not, at `runs/<task-type>/worker-results/<worker>-audit-<task-type>-<seq>.md`. Workers write both in the same step, so a result without a sidecar means the Reading Confirmation block — the only evidence the worker read its inputs — was never produced. `validate-run.py` fails the run on this at Phase 7 either way (`validate_worker_results_audit`); checking it here spends the existing one-retry budget while the role can still be re-dispatched, instead of surfacing hours later when the worker session is gone.
|
|
172
|
-
- `okstra worker-audit-check --run-dir <
|
|
174
|
+
- `okstra worker-audit-check --run-dir <runDirectoryPath from the active run context> --task-type <t> --seq <n> --worker <id>` exits 2 on a backticked `path:line` citation in the worker's result that has no matching Evidence read row in its audit sidecar. Run it the moment you collect each result. Phase 7 enforces the same rules from the same implementation (`okstra_ctl.worker_audit_ledger`), but by then the worker session is gone and the only remaining moves are editing the result yourself — which destroys the audit chain the ledger exists to provide — or ending the run `contract-violated`. While the session is alive, `SendMessage` to the worker so it corrects its own citation; that costs about a minute against a re-dispatch or a failed run.
|
|
173
175
|
|
|
174
176
|
The same audit check parses command evidence from canonical rows such as `- Evidence command: {"command":"npm run check","cwd":"<project-root>","exitCode":0,"outputSummary":"all checks passed"}`. Workers record only commands that produced or verified a conclusion, not exploratory `rg`, `ls`, or file-opening commands. A malformed row is a contract failure returned by `okstra_ctl.worker_audit_ledger`; send the failure to the worker while its session is still available. Nothing scans these rows for secrets, so a row is never refused for how a setting is named — `AUTH_MODE=off` and `TOKEN_TTL=60` record verbatim. Do not tell a worker to drop or paraphrase a row on those grounds: a worker that cannot record the command it ran has no way to satisfy the evidence ledger that demands it.
|
|
175
177
|
|
|
176
178
|
**One-retry policy:**
|
|
177
179
|
|
|
178
|
-
1. On the FIRST result-missing trigger for a
|
|
180
|
+
1. On the FIRST result-missing trigger for a role within a run, use the adapter's retry operation with the same logical instructions, Result Path, model assignment, and role identity. V2 retries use `materialize_retry_invocation`: it preserves the previous prompt and metadata, generates the next attempt's Prompt History Path, metadata path, and attempt number, and verifies the new pair. The logical `invocationRef` remains unchanged. Do not copy the prior prompt's attempt fields or create a new role slot.
|
|
179
181
|
2. If the SECOND attempt also fails the same check, Lead records the role's terminal status as `error` with `--message "result-missing after 1 retry"` and proceeds to Phase 5.5 / Phase 6 with the remaining workers' results. Lead MUST NOT retry a third time — convergence and the report writer are designed to operate on reduced-confidence single-or-two-analyser mode when one role is absent (`prompts/lead/okstra-lead-contract.md` "If only one worker result is usable: reduced-confidence synthesis").
|
|
180
182
|
3. The retry counter is **per-run, per-role** and is NOT preserved across runs. A subsequent okstra run for the same task-key starts each role's counter fresh.
|
|
181
183
|
4. Convergence reverify rounds (Phase 5.5) inherit the same one-retry budget — a reverify dispatch that triggers result-missing may be re-dispatched once.
|
|
182
|
-
5. **Reading-plan exception for a deterministic failure.**
|
|
183
|
-
- The exception applies only when the first attempt produced **no result file** after consuming its full timeout. A worker that failed fast, errored, or returned a partial artifact
|
|
184
|
+
5. **Reading-plan exception for a deterministic failure.** When the first attempt exhausted its time budget without a result and measured input size explains that failure, the lead may revise the reading plan through a fresh corrective invocation. Preserve the Result Path, model assignment, and role identity. Use `agent-prompt materialize` with a fresh invocation ID and prompt path; changed task instructions cannot reuse a logical invocation whose input digest is already reserved. The new prompt supplies its own execution headers.
|
|
185
|
+
- The exception applies only when the first attempt produced **no result file** after consuming its full timeout. A worker that failed fast, errored, or returned a partial artifact follows the ordinary retry contract above.
|
|
184
186
|
- Lead MUST log the deviation with the normal `contract-deviation` entry naming the removed or reordered reads and the measured evidence that motivated it (byte counts, elapsed time). An unlogged reading-plan change is a contract violation, not an exception.
|
|
185
187
|
- This exception never licenses changing what the worker is asked to *produce*. Narrowing the deliverable to make it finish is a contract violation.
|
|
186
188
|
|
|
@@ -199,6 +201,8 @@ Lead-facing duties that stay here:
|
|
|
199
201
|
- Lead and report-writer MUST preserve worker ticket tagging and `worker:item` source pairs when carrying findings into the final report.
|
|
200
202
|
- Lead's result-missing redispatch check parses the frontmatter and sections 1–5 presence per the preamble contract (see "Lead Redispatch Policy on Result-Missing").
|
|
201
203
|
|
|
204
|
+
**Enforced:** `scripts/okstra_ctl/convergence.py` `_parse_group_source` rejects a finding source that is not a `worker:item-id` pair, so a finding stripped of its worker attribution cannot enter the convergence queue the report is assembled from.
|
|
205
|
+
|
|
202
206
|
### Worker error-log command
|
|
203
207
|
|
|
204
208
|
Workers record a tool failure through the typed `okstra error-log append-observed` command in `templates/worker-error-contract.md`. They do not create a JSON sidecar or pass JSON text. The command receives the assigned worker identity, task key, phase, model, and observed scalar fields; the Python writer owns JSONL serialization and validation.
|
|
@@ -215,6 +219,8 @@ Workers record a tool failure through the typed `okstra error-log append-observe
|
|
|
215
219
|
- **No external timeout around `worker-dispatch`.** The deterministic dispatch owns BOTH timeout mechanisms: (1) the process polling cap (30min + optional 5min mtime grace), and (2) the shared runner's stream-idle watchdog. If the CLI produces nothing for `<idle-timeout-seconds>`, the runner terminates the process group and marks the status sidecar timed out. Lead MUST NOT layer an earlier host timeout around it.
|
|
216
220
|
- `contract-violation` events (C) are recorded by Lead via `okstra error-log append-observed --error-type contract-violation ...` after inspecting worker outputs.
|
|
217
221
|
|
|
222
|
+
**Enforced:** `scripts/okstra_ctl/worker_prompt_policy.py` `ERRORS_PATH_HEADERS` is the literal the prompt emitter writes and `scripts/okstra_ctl/worker_prompt_contract.py` treats as a non-body header, so the anchor reaches the worker rather than being synthesized; `scripts/okstra_ctl/dispatch_core.py` owns the polling cap and the runner's idle watchdog, leaving no place for a lead-side timeout.
|
|
223
|
+
|
|
218
224
|
## Convergence Phase Rules
|
|
219
225
|
|
|
220
226
|
1. Re-verification uses the same worker roles and model assignments as the initial run.
|
|
@@ -44,6 +44,7 @@ profile document.
|
|
|
44
44
|
`[CONFIRMED <YYYY-MM-DD> → RC-N]` markers on `Open Questions` rows are the per-row signal that the reporter has answered; their answers live verbatim under `## Reporter Confirmations` in the brief.
|
|
45
45
|
- `Source Material` is reporter-verbatim. Do NOT paraphrase, summarize, reorder, or restructure it. Quote it directly when needed.
|
|
46
46
|
- `## Task-Group Context` (present in the analysis packet only when the task-group has `.okstra/briefs/<task-group>/group-context.md`, copied to `instruction-set/task-group-context.md`) is background shared by every task in the group — why the group exists, how the result is measured, group-wide constraints, ticket relations. It is not this task's requirement ledger: the brief decides this run's scope, the group context decides the measure and what no task may break. Reason from the group's measure (its stated denominator) rather than from a quantity only this ticket exposes, and where the brief and the group context conflict, raise a `Clarification Items` row instead of resolving the conflict silently in either direction.
|
|
47
|
+
- `## Task-Group Memory` (present after `## Task-Specific Brief Extract` when a sibling task of the group has recorded a run) is what the other tasks of this task-group concluded in their latest run: per task, the phase it reached and its next phase, the headline, accumulated decisions, watch-out items, open follow-ups, and the record path. okstra writes it into the group's `group-context.md` at every `report-finalize` (`record-group-memory`); the region also opens with the group's start order (brief ordinal; graph edges quoted as `waits for`). A headline is a claim — read the cited record before building on it. A sibling's decision does not bind this run: where it conflicts with the brief, raise a `Clarification Items` row rather than re-deriving what the sibling settled or overriding it silently. This task's own entry is not carried; its history arrives through the timeline, prior-planning, and fix-history sections.
|
|
47
48
|
- `Related Task Graph` is the structured task-topology handoff. If the section is present and not `_(none)_`, read it before classification, diagnosis, candidate discovery, fan-out, or next-step routing. Preserve the edge direction exactly as written: `From` → `To` is load-bearing for `depends-on`, `blocks`, parent/child, follow-up, and split relations.
|
|
48
49
|
- Valid relations are `parent-of`, `child-of`, `depends-on`, `blocks`, `blocked-by`, `follow-up-of`, `split-from`, `duplicates`, and `related-to`. `duplicates` / `related-to` are undirected; every other relation is directed.
|
|
49
50
|
- Treat `Source` as the edge provenance. Do not invent graph edges from filename similarity, topic overlap, or worker preference. If a needed edge is missing or contradictory, raise a `Clarification Items` row instead of silently reordering tasks.
|
|
@@ -90,7 +91,7 @@ profile document.
|
|
|
90
91
|
- if a schema-v1 table or an analysis-worker result table requires a recommended answer, alternatives, or an evidence-check note, encode it inside the existing 4-column schema: put evidence notes in `Statement` as `Evidence checked: <path:line>` or `Evidence checked: none — <human-only reason>`, and put recommendations/options in `Expected form` as `Recommended: (a) <answer> — <rationale>; Alternatives: (b) <option> (c) <option>`. The recommended answer is always the first option and MUST carry the `(a)` label; alternatives continue the same letter sequence from `(b)` (a lone alternative is `(b) <option>`, never restart at `(a)`), so the full option set reads `(a) (b) (c) …` in order and renders each as its own selectable option. Do **not** append a pick-one answer-space summary such as `(pick 1 of A / B)` or `(pick N of …)` to `<options>` — the rendered `<select>` already enforces single choice, and that annotation leaks verbatim into an option label. Do not add `Recommended`, `Evidence`, `Alternatives`, or `evidence-checked` columns, and do not break the merged record-meta cell back into separate columns.
|
|
91
92
|
- For schema v2, data.json is canonical and the HTML exports answers to a user-response sidecar; the source report is never edited. `--resume-clarification` carries those answers into the next run. The lower-level `--clarification-response <path>` remains available for scripted runs.
|
|
92
93
|
- When a response is carried in, reconcile every prior `clarificationItems[]` row against new evidence and update its status to `resolved` or `obsolete` before issuing the next verdict. Schema-v1 compatibility Markdown may additionally render its conditional Section 0; the schema-v2 full reading copy records decisions under `## Clarification and User Decisions`.
|
|
93
|
-
- **Supersession (BLOCKING).** Reconciling the `C-*` row is only half of incorporating an answer. An answer does not merely *add* a decision — it *invalidates* whatever the previous run wrote under the opposite assumption. Before issuing the next decision, walk the prior deliverable prose for every statement the answer makes false and **delete or rewrite it**, then record the retirement. Adding the new decision while leaving the contradicting sentence in place puts two opposite instructions for the same symbol in one document; the implementer must then guess which is live, and the next verification round correctly blocks on it. In `implementation-planning` this record is `implementationPlanning.supersessionLedger[]` — one entry per answered clarification, either `disposition: superseded` (with the retired statement, its replacement, and the sections revised) or `disposition: no-dependent-statement` (with a rationale). **Enforced:** `validators/validate-run.py` `_validate_supersession_ledger` requires an entry per answered clarification; whether the claim is *true* is what the §5.5.9 adversarial round tests.
|
|
94
|
+
- **Supersession (BLOCKING).** Reconciling the `C-*` row is only half of incorporating an answer. An answer does not merely *add* a decision — it *invalidates* whatever the previous run wrote under the opposite assumption. Before issuing the next decision, walk the prior deliverable prose for every statement the answer makes false and **delete or rewrite it**, then record the retirement. Adding the new decision while leaving the contradicting sentence in place puts two opposite instructions for the same symbol in one document; the implementer must then guess which is live, and the next verification round correctly blocks on it. In `implementation-planning` this record is `implementationPlanning.supersessionLedger[]` — one entry per answered clarification, either `disposition: superseded` (with the retired statement, its replacement, and the sections revised) or `disposition: no-dependent-statement` (with a rationale). A plan built from a selected direction inherits the answers the option-selection record carried before it has any statement to retire, so those carried rows need no entry; the rows this plan itself raised and settled still do. **Enforced:** `validators/validate-run.py` `_validate_supersession_ledger` requires an entry per answered clarification, exempting the ledger's `carriedDecisions[]` ids on a selected-direction plan; whether the claim is *true* is what the §5.5.9 adversarial round tests.
|
|
94
95
|
- Verdict Card data consistency (shared; schema-v1 Markdown keeps the legacy visible card):
|
|
95
96
|
- The Card carries no verdict token — the token lives once, in `finalVerdict.verdictToken`, and every gate reads it there. `verdictCard.direction` byte-matches `finalVerdict.direction`; next-step routing agrees with `recommendedNextSteps[0]`. The full reading copy and human summary are derived from the data fields without repeating both visible sections. **Enforced in part:** the v3.0 schema's `verdictCard` is `additionalProperties: false` with no verdict-token property, so the token cannot be duplicated onto the Card, and `scripts/okstra_ctl/report_narrative.py` `_writer_owned_schema` applies the finished report's `$defs.Direction` enum to the narrative, rejecting an off-enum `direction` while the writer can still be re-run. The byte-match between the two `direction` fields is not compared by anything — assembly overwrites `nextStep` on both when the plan-body gate passes (`scripts/okstra_ctl/report_assembly.py:590-604`) but leaves `direction` as the writer wrote it.
|
|
96
97
|
- Cross-worker traceability (shared — applies to every analysis worker output and to the lead's `## 6.` / `## 2.` tables in the final-report):
|