okstra 0.168.0 → 0.169.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +5 -4
- package/docs/architecture/storage-model.md +57 -1
- package/docs/architecture.md +70 -2
- package/docs/cli.md +8 -4
- package/docs/for-ai/skills/okstra-code-review.md +3 -2
- package/docs/for-ai/skills/okstra-schedule-gen.md +3 -1
- package/docs/project-structure-overview.md +14 -11
- package/package.json +1 -1
- package/runtime/BUILD.json +2 -2
- package/runtime/agents/workers/claude-worker.md +6 -5
- package/runtime/agents/workers/report-writer-worker.md +9 -4
- package/runtime/agents/workers/translator-worker.md +6 -4
- package/runtime/bin/okstra-error-log.py +38 -282
- package/runtime/prompts/duties/acceptance-critic.md +24 -0
- package/runtime/prompts/duties/acceptance-verifier.md +24 -0
- package/runtime/prompts/duties/analysis-worker.md +24 -0
- package/runtime/prompts/duties/code-reviewer.md +24 -0
- package/runtime/prompts/duties/common.md +35 -0
- package/runtime/prompts/duties/implementation-executor.md +24 -0
- package/runtime/prompts/duties/implementation-verifier.md +24 -0
- package/runtime/prompts/duties/lead.md +24 -0
- package/runtime/prompts/duties/report-writer.md +24 -0
- package/runtime/prompts/duties/reverification-worker.md +24 -0
- package/runtime/prompts/duties/schedule-verifier.md +24 -0
- package/runtime/prompts/duties/scope-critic.md +24 -0
- package/runtime/prompts/duties/translator.md +24 -0
- package/runtime/prompts/lead/convergence.md +104 -14
- package/runtime/prompts/lead/okstra-lead-contract.md +11 -21
- package/runtime/prompts/lead/plan-body-verification.md +16 -1
- package/runtime/prompts/lead/report-writer.md +20 -5
- package/runtime/prompts/lead/team-contract.md +13 -13
- package/runtime/prompts/profiles/_coding-conventions-preflight.md +1 -1
- package/runtime/prompts/profiles/_implementation-diff-review.md +1 -1
- package/runtime/prompts/profiles/_implementation-executor.md +1 -1
- package/runtime/prompts/profiles/implementation.md +4 -2
- package/runtime/python/okstra_ctl/adapters/hosts/antigravity/adapter.py +6 -0
- package/runtime/python/okstra_ctl/adapters/hosts/antigravity/relay.md +3 -2
- package/runtime/python/okstra_ctl/adapters/hosts/capability_adapter.py +8 -0
- package/runtime/python/okstra_ctl/adapters/hosts/claude-code/adapter.py +33 -0
- package/runtime/python/okstra_ctl/adapters/hosts/claude-code/relay.md +13 -12
- package/runtime/python/okstra_ctl/adapters/hosts/codex/adapter.py +6 -0
- package/runtime/python/okstra_ctl/adapters/hosts/codex/relay.md +3 -2
- package/runtime/python/okstra_ctl/adapters/hosts/external/adapter.py +2 -0
- package/runtime/python/okstra_ctl/adapters/hosts/external/relay.md +3 -3
- package/runtime/python/okstra_ctl/adapters/hosts/grok/adapter.py +6 -0
- package/runtime/python/okstra_ctl/adapters/hosts/grok/relay.md +2 -1
- package/runtime/python/okstra_ctl/adapters/hosts/kimi/adapter.py +6 -0
- package/runtime/python/okstra_ctl/adapters/hosts/kimi/relay.md +2 -1
- package/runtime/python/okstra_ctl/agent_invocation.py +1582 -0
- package/runtime/python/okstra_ctl/agent_prompt_cli.py +796 -0
- package/runtime/python/okstra_ctl/codex_dispatch.py +2 -107
- package/runtime/python/okstra_ctl/context_cost.py +46 -5
- package/runtime/python/okstra_ctl/dispatch_core.py +538 -43
- package/runtime/python/okstra_ctl/dispatch_state.py +461 -36
- package/runtime/python/okstra_ctl/doctor.py +90 -16
- package/runtime/python/okstra_ctl/entrypoints/hosts.py +87 -9
- package/runtime/python/okstra_ctl/error_log_write.py +308 -0
- package/runtime/python/okstra_ctl/initial_prompt_materialization.py +214 -23
- package/runtime/python/okstra_ctl/path_hints.py +26 -0
- package/runtime/python/okstra_ctl/paths.py +20 -0
- package/runtime/python/okstra_ctl/ports/__init__.py +8 -0
- package/runtime/python/okstra_ctl/ports/host.py +3 -0
- package/runtime/python/okstra_ctl/ports/host_model.py +60 -0
- package/runtime/python/okstra_ctl/registry/host_registry.py +5 -0
- package/runtime/python/okstra_ctl/render.py +217 -12
- package/runtime/python/okstra_ctl/report_finalize.py +44 -0
- package/runtime/python/okstra_ctl/run.py +368 -51
- package/runtime/python/okstra_ctl/session.py +16 -12
- package/runtime/python/okstra_ctl/team.py +11 -11
- package/runtime/python/okstra_ctl/worker_audit_check.py +26 -4
- package/runtime/python/okstra_ctl/worker_audit_ledger.py +59 -9
- package/runtime/python/okstra_ctl/worker_dispatch.py +104 -0
- package/runtime/python/okstra_ctl/worker_prompt_body.py +5 -38
- package/runtime/python/okstra_ctl/worker_prompt_contract.py +56 -3
- package/runtime/python/okstra_ctl/worker_prompt_headers.py +2 -2
- package/runtime/python/okstra_ctl/worker_prompt_policy.py +38 -1
- package/runtime/skills/okstra-code-review/SKILL.md +22 -3
- package/runtime/skills/okstra-run/SKILL.md +16 -1
- package/runtime/skills/okstra-schedule-gen/SKILL.md +15 -1
- package/runtime/templates/implementation-worker-preamble.md +0 -10
- package/runtime/templates/report-writer-prompt-preamble.md +0 -9
- package/runtime/templates/reports/settings.template.json +0 -11
- package/runtime/templates/worker-prompt-preamble.md +0 -10
- package/runtime/validators/lib/fixtures.sh +93 -0
- package/runtime/validators/lib/validate-assets.sh +0 -8
- package/runtime/validators/validate-run.py +182 -0
- package/src/cli-registry.mjs +14 -0
- package/src/commands/execute/agent-prompt.mjs +25 -0
- package/src/commands/execute/codex-dispatch.mjs +6 -63
- package/src/commands/execute/worker-dispatch.mjs +76 -0
- package/src/commands/lifecycle/doctor.mjs +18 -3
- package/src/commands/lifecycle/install.mjs +33 -15
- package/src/commands/lifecycle/uninstall.mjs +4 -3
- package/src/lib/install-assets.mjs +9 -0
- package/runtime/agents/workers/antigravity-worker.md +0 -259
- package/runtime/agents/workers/codex-worker.md +0 -259
- package/runtime/agents/workers/grok-worker.md +0 -259
- package/runtime/agents/workers/kimi-worker.md +0 -259
- package/runtime/prompts/coding-preflight/scripts/preedit-check.sh +0 -79
- package/runtime/templates/operating-standard.md +0 -22
- package/src/lib/worker-agent-render.mjs +0 -50
|
@@ -224,6 +224,40 @@ Design intent: one `counter-evidence` refute denies a claim consensus (it cannot
|
|
|
224
224
|
|
|
225
225
|
## Re-verification Dispatch
|
|
226
226
|
|
|
227
|
+
### Invocation materialization gate (BLOCKING)
|
|
228
|
+
|
|
229
|
+
For every finding reverify row and critic-gap verification row, first write a
|
|
230
|
+
call-specific task-instructions file under the current run's `state/`
|
|
231
|
+
directory. Then run `okstra agent-prompt materialize` with `--audience
|
|
232
|
+
reverification-worker`, `--assignment-ref reverify/<workerId>`, the exact
|
|
233
|
+
`--worker-id`, `--dispatch-kind reverify-r<N>`, and the authorized
|
|
234
|
+
prompt/result/audit paths. The returned `promptPath` is the only body that may
|
|
235
|
+
be dispatched; do not append role prose or reconstruct model headers after
|
|
236
|
+
materialization.
|
|
237
|
+
|
|
238
|
+
If the dispatch gate then rejects that prompt, fix the task-instructions file
|
|
239
|
+
and re-run the same `materialize` call with `--replace-undispatched` — keep the
|
|
240
|
+
`--invocation-id`. A published prompt is otherwise immutable, so without that
|
|
241
|
+
flag the retry fails as `existing_invocation_conflict`; do NOT delete the
|
|
242
|
+
reservation under `prompts/.agent-invocations` and do NOT mint a second
|
|
243
|
+
invocation id to get around it, because both detach the audit chain from the
|
|
244
|
+
call it describes. The flag is checked: once any dispatch row names this
|
|
245
|
+
invocation, the prompt is history and the replacement is refused.
|
|
246
|
+
|
|
247
|
+
Run `okstra agent-prompt verify --run-manifest <path> --metadata
|
|
248
|
+
<metadataPath> --json` immediately before dispatch. A failed verification is a
|
|
249
|
+
pre-dispatch contract failure. For `runner=native-session`, pass only the
|
|
250
|
+
returned `hostModelValue` to the host model argument. For
|
|
251
|
+
`runner=cli-wrapper`, invoke `okstra worker-dispatch` and let it consume the
|
|
252
|
+
returned `modelExecutionValue`; never pass that value as a native-host model
|
|
253
|
+
token. Before a host-native call, run `okstra agent-prompt record-dispatch`
|
|
254
|
+
with the project root, run manifest, metadata path, and
|
|
255
|
+
`--enforcement-mode host-native-spec-link-gate`. After its Result Path exists,
|
|
256
|
+
run `okstra agent-prompt link-result` with the same run manifest,
|
|
257
|
+
`--dispatch-id <invocationId>:attempt-1`, and that result path before reading
|
|
258
|
+
the result. This link proves that the accepted result belongs to a verified
|
|
259
|
+
call specification; it does not prove which bytes the host primitive delivered.
|
|
260
|
+
|
|
227
261
|
### Sponsorship Optimization
|
|
228
262
|
|
|
229
263
|
For each persisted round plan, build exactly one prompt per `dispatches[]` row and call `redispatch_worker(assignment, prompt, reason)` once through the selected runtime adapter. The prompt contains exactly that row's `findingIds` in plan order and MUST NOT add, remove, or reorder findings. This excludes Section 6, every resolved finding, and every finding owned by the receiving origin worker because none can appear in the engine row. The assignment, model, prompt path, Result Path, worker-results path, errors paths, and `dispatchKind` come from the current run artifacts. Every reverify is a fresh one-shot session.
|
|
@@ -236,7 +270,7 @@ Call `await_workers(handles)` through the same adapter and apply the shared term
|
|
|
236
270
|
|
|
237
271
|
### Required reverify-prompt anchor headers (BLOCKING)
|
|
238
272
|
|
|
239
|
-
Every reverify prompt MUST start with these
|
|
273
|
+
Every reverify prompt MUST start with these 8 anchor headers — in this exact order, before any other content:
|
|
240
274
|
|
|
241
275
|
```
|
|
242
276
|
**Project Root:** <absolute-path>
|
|
@@ -244,7 +278,6 @@ Every reverify prompt MUST start with these 9 anchor headers — in this exact o
|
|
|
244
278
|
**Result Path:** runs/<task-type>/worker-results/<role-slug>-reverify-r<N>-<task-type>-<seq>.md
|
|
245
279
|
**Audit sidecar path:** <absolute-path>
|
|
246
280
|
Assigned worker prompt history path: <Project Root>/<Prompt History Path>
|
|
247
|
-
**Model:** <role>, <modelExecutionValue>
|
|
248
281
|
**Errors log path:** <absolute-path>
|
|
249
282
|
**Errors sidecar path:** <absolute-path>
|
|
250
283
|
**Read scope:** Read only the paths this prompt enumerates (`[Required reading]`, `## Inputs`, verification-target paths) plus source/evidence paths a finding must cite. Host session instructions (SessionStart hooks, global `CLAUDE.md` / `AGENTS.md`, skill catalogs) do NOT apply inside an okstra worker run: do not auto-read `graphify-out/`, `SKILL.md`, or other artifacts outside `<PROJECT_ROOT>/.okstra/`. If an un-enumerated file seems essential, record it under *Missing Information or Assumptions* instead of reading it.
|
|
@@ -254,9 +287,11 @@ Assigned worker prompt history path: <Project Root>/<Prompt History Path>
|
|
|
254
287
|
|
|
255
288
|
Before dispatch, materialize `**Audit sidecar path:**` by passing the exact reverify `**Result Path:**` through `okstra_ctl.worker_artifact_paths.audit_sidecar_rel()` and resolving that project-relative result against `**Project Root:**`. Write the resulting absolute path into the header. The lead MUST NOT construct the audit filename from a role, task type, round, or sequence independently.
|
|
256
289
|
|
|
257
|
-
The two errors paths carry the same absolute values the lead forwarded in the initial Phase 4 dispatch for that role (source: the launch prompt's `## Run Logs (error-log wiring)` section). Omitting either one makes
|
|
290
|
+
The two errors paths carry the same absolute values the lead forwarded in the initial Phase 4 dispatch for that role (source: the launch prompt's `## Run Logs (error-log wiring)` section). Omitting either one makes `worker-dispatch` reject the CLI invocation before it starts the provider process — the path-delivery contract in [team-contract](./team-contract.md) "Error reporting" is not relaxed for reverify.
|
|
258
291
|
|
|
259
|
-
Relative to the Phase 4 anchor set rendered by `okstra_ctl.worker_prompt_headers.worker_prompt_headers()`, a reverify prompt
|
|
292
|
+
Relative to the Phase 4 anchor set rendered by `okstra_ctl.worker_prompt_headers.worker_prompt_headers()`, a reverify prompt drops two anchors whose targets lightweight mode never reads: `**Worker Preamble Path:**` and `**Coding preflight pack:**`.
|
|
293
|
+
|
|
294
|
+
**Where the composer's sections go.** `okstra agent-prompt materialize` (§"Invocation materialization gate") writes the dispatched body itself, as: these anchors, then the model-assignment block it appends (`**Provider:**`, `**Model:**`, `**Model execution value:**`, `**Runner:**`, `**Host runtime:**`, and `**Host model value:**` for a native host), then `## Duty Contract`, then `## Task Instructions` followed verbatim by the task-instructions file the lead wrote. So the lead authors only the last part, and every rule below about ordering — the phase boundary before the instruction headings, the `**Model:** <role>, <modelExecutionValue>` line — is about the lead's own file, not about the composed document. The composer's `**Model:** <modelExecutionValue>` anchor is a different line with a different shape; do not try to reshape it, and do not count it among the 8.
|
|
260
295
|
|
|
261
296
|
For an `antigravity` assignment, append the exact `PLAIN_FILE_WRITE_HEADER`
|
|
262
297
|
value from `okstra_ctl.worker_prompt_headers` immediately after
|
|
@@ -265,20 +300,28 @@ persisted initial Phase 4 prompt; do not paraphrase or reconstruct it. If the
|
|
|
265
300
|
persisted initial prompt does not contain that generated header, abort the
|
|
266
301
|
reverify dispatch and record a `contract-violation` event instead of
|
|
267
302
|
dispatching without the plain-file safeguard. This provider-specific header is
|
|
268
|
-
outside the common
|
|
303
|
+
outside the common 8-header count above. Other providers do not receive it.
|
|
269
304
|
|
|
270
305
|
The rationale for both drops is §"Reverify prompt: required-reading suppression" below.
|
|
271
306
|
|
|
272
|
-
|
|
273
|
-
|
|
274
|
-
instructions:
|
|
307
|
+
The task-instructions file the lead writes MUST open with this block, before any
|
|
308
|
+
`##` heading of its own:
|
|
275
309
|
|
|
276
310
|
```markdown
|
|
311
|
+
**Model:** <role>, <modelExecutionValue>
|
|
277
312
|
**Task Type:** <task-manifest taskType>
|
|
278
313
|
**Forbidden actions:**
|
|
279
314
|
<active-run-context workflow.forbiddenActions, verbatim>
|
|
280
315
|
```
|
|
281
316
|
|
|
317
|
+
This is the same placement `okstra_ctl.worker_prompt_body` uses for an initial
|
|
318
|
+
Phase 4 prompt, and it is where the checks look: `validate_reverify_prompt()`
|
|
319
|
+
reads the region after `## Task Instructions`, so a `**Model:**` line left in
|
|
320
|
+
the anchors is invisible to it, and a phase boundary written below the file's
|
|
321
|
+
first heading fails. Both rules are about this file — the composer's own
|
|
322
|
+
`## Duty Contract` heading sits above everything here and is not what they
|
|
323
|
+
measure against.
|
|
324
|
+
|
|
282
325
|
Do not summarize, shorten, or reconstruct the forbidden-actions text. The
|
|
283
326
|
selected adapter validates the task type and exact block through
|
|
284
327
|
`okstra_ctl.worker_prompt_contract.validate_reverify_prompt()` before starting
|
|
@@ -323,7 +366,7 @@ This is the single largest avoidable cost in `requirements-discovery`, `error-an
|
|
|
323
366
|
### Lightweight Re-verification Prompt
|
|
324
367
|
|
|
325
368
|
```
|
|
326
|
-
|
|
369
|
+
Perform re-verification for <task-key> (round <N>).
|
|
327
370
|
|
|
328
371
|
## Instructions
|
|
329
372
|
|
|
@@ -365,7 +408,7 @@ For each finding, respond as:
|
|
|
365
408
|
Used instead of the lightweight/full-reanalysis prompt when `config.adversarial == true`. The required anchor headers (§"Required reverify-prompt anchor headers") are identical. The `[Required reading]` clause is suppressed; only the cited-evidence paths of the items under attack are injected (see §"Adversarial Verification Mode" → Scoped full-reanalysis).
|
|
366
409
|
|
|
367
410
|
```
|
|
368
|
-
|
|
411
|
+
Perform ADVERSARIAL re-verification for <task-key> (round <N>).
|
|
369
412
|
|
|
370
413
|
## Instructions
|
|
371
414
|
|
|
@@ -416,7 +459,7 @@ UNVERIFIABLE is **not** `verification-error`. A verifier that opened the evidenc
|
|
|
416
459
|
### Full Re-analysis Re-verification Prompt
|
|
417
460
|
|
|
418
461
|
```
|
|
419
|
-
|
|
462
|
+
Perform deep re-verification for <task-key> (round <N>).
|
|
420
463
|
|
|
421
464
|
## Instructions
|
|
422
465
|
|
|
@@ -550,7 +593,45 @@ The critic input is the Round 0 consolidated finding list. Reverify rounds only
|
|
|
550
593
|
- **Gap verification + merge**: only after BOTH the finding-convergence loop has exited AND the critic result is collected, and BEFORE the Phase 6 report-writer dispatch. If the loop exited `aborted-non-result`, do NOT dispatch a gap-verification round — record every gap in `unverifiedGaps[]` per §"Gap verification".
|
|
551
594
|
|
|
552
595
|
### Dispatch (fresh one-shot)
|
|
553
|
-
|
|
596
|
+
Write the critic-only task instructions, then run `okstra agent-prompt
|
|
597
|
+
materialize` with `--audience scope-critic`, `--assignment-ref critic/scope`,
|
|
598
|
+
the critic worker ID, and `--dispatch-kind critic`. Verify the returned
|
|
599
|
+
`metadataPath` before dispatch and use its `promptPath` without modification.
|
|
600
|
+
For `runner=native-session`, use only `hostModelValue`; for
|
|
601
|
+
`runner=cli-wrapper`, use `okstra worker-dispatch`, which consumes
|
|
602
|
+
`modelExecutionValue`. Record host-native linkage with
|
|
603
|
+
`enforcementMode=host-native-spec-link-gate` and the metadata path. If the
|
|
604
|
+
persisted assignment or either model value required by its runner is absent,
|
|
605
|
+
record `critic-skipped: model-unresolved`; never resolve a replacement model.
|
|
606
|
+
Result path: `runs/<task-type>/worker-results/<provider>-worker-critic-<task-type>-<seq>.md`.
|
|
607
|
+
|
|
608
|
+
**What the critic task-instructions file MUST contain (BLOCKING).** A critic
|
|
609
|
+
dispatch is not a reverify dispatch: `dispatchKind = "critic"` keeps
|
|
610
|
+
`audience = "analysis"`, so `worker_prompt_contract.validate_initial_prompts`
|
|
611
|
+
judges it by the full initial-analysis contract. Two of those requirements are
|
|
612
|
+
satisfied by the generated body for a Phase 4 worker
|
|
613
|
+
(`okstra_ctl.worker_prompt_body`) and by nothing at all for a critic, whose
|
|
614
|
+
instructions the lead writes — the materializer's anchor block supplies neither.
|
|
615
|
+
Put both in the instructions file:
|
|
616
|
+
|
|
617
|
+
```markdown
|
|
618
|
+
**Prompt Delivery Mode:** eager-include
|
|
619
|
+
```
|
|
620
|
+
|
|
621
|
+
and, under the file's `## Inputs`, exactly one line in this shape — the literal
|
|
622
|
+
label and the backticks are what the check matches, so a bare path or a
|
|
623
|
+
differently-worded label counts as zero:
|
|
624
|
+
|
|
625
|
+
```markdown
|
|
626
|
+
- Primary analysis packet: `<path ending in analysis-packet.md>`
|
|
627
|
+
```
|
|
628
|
+
|
|
629
|
+
Omitting either one fails `okstra team dispatch --dispatch-kind critic` before
|
|
630
|
+
any process starts, reported as `<task-type> prompt contract: <worker>: exactly
|
|
631
|
+
one Primary analysis packet path is required (found 0)` and `exactly one
|
|
632
|
+
non-empty **Prompt Delivery Mode:** header is required`. Fix the instructions
|
|
633
|
+
file and re-materialize with `--replace-undispatched` (§"Invocation
|
|
634
|
+
materialization gate") rather than editing the published prompt.
|
|
554
635
|
|
|
555
636
|
The `-worker-` token is load-bearing, not decoration: the critic prompt carries the same generated anchor headers as every other worker ([team-contract](./team-contract.md) §"Worker prompts"), and its `**Audit sidecar path:**` comes from passing that result path through `okstra_ctl.worker_artifact_paths.audit_sidecar_rel()`, which inserts `-audit-` after the token and raises without it. A `<provider>-critic-...` name leaves the lead choosing between breaking the contract and hand-inventing the sidecar name. Note that `originWorker` stays `"<provider>-critic"` — that is a worker id in the convergence state, not a filename, and the two do not have to match.
|
|
556
637
|
|
|
@@ -567,7 +648,7 @@ Required reading before proposing a gap or an over-scope candidate:
|
|
|
567
648
|
Operational guardrails are not task requirements. A gap must trace to a brief requirement, an analysis-packet scope item, a source path the packet authorizes, or an evidence claim in a worker result. Do NOT infer missing verification from a one-line summary; open the named result and audit sidecar first.
|
|
568
649
|
|
|
569
650
|
```
|
|
570
|
-
|
|
651
|
+
Inspect scope coverage for <task-key>. Below are the consolidated findings the
|
|
571
652
|
workers produced. Your job has exactly two halves. Answer both.
|
|
572
653
|
|
|
573
654
|
(1) MISSING — name what nobody covered:
|
|
@@ -617,10 +698,19 @@ The asymmetry is deliberate and runs the opposite way from the coverage half: a
|
|
|
617
698
|
|
|
618
699
|
The `final-verification` phase uses the same fresh one-shot `redispatch_worker` pattern and the same dispatch timing as §"Coverage critic pass" §"When" (provider + `config.critic.modelExecutionValue` from the `convergence.critic` block; default off; same model-unresolved skip rule) — the delivered work the critic inspects is likewise fixed before the reverify round starts. Only the prompt, the verification semantics, and the output sink differ — final-verification's findings are defects/blockers, so the critic acts as an **acceptance devil's advocate** (find reasons NOT to accept), and its candidate blockers are NEVER dropped (that would suppress real defects).
|
|
619
700
|
|
|
701
|
+
Before that call, write the acceptance-only task instructions and run `okstra
|
|
702
|
+
agent-prompt materialize` with `--audience acceptance-critic`,
|
|
703
|
+
`--assignment-ref critic/acceptance`, the critic worker ID, and
|
|
704
|
+
`--dispatch-kind critic`. Verify the returned `metadataPath`, dispatch only the
|
|
705
|
+
returned `promptPath`, and select `hostModelValue` for a native host or
|
|
706
|
+
`modelExecutionValue` through `okstra worker-dispatch`. Native dispatch linkage
|
|
707
|
+
uses `enforcementMode=host-native-spec-link-gate`; it does not claim prompt
|
|
708
|
+
delivery was observed.
|
|
709
|
+
|
|
620
710
|
### Prompt
|
|
621
711
|
|
|
622
712
|
```
|
|
623
|
-
|
|
713
|
+
Challenge acceptance for <task-key>. The delivered work is about
|
|
624
714
|
to be judged for acceptance. Your ONLY job is to find reasons it should NOT be
|
|
625
715
|
accepted — surface candidate acceptance BLOCKERS the verifiers may have missed:
|
|
626
716
|
- requirements / acceptance points with no covering evidence,
|
|
@@ -1,15 +1,5 @@
|
|
|
1
1
|
# Okstra Lead Contract
|
|
2
2
|
|
|
3
|
-
## Operating standard
|
|
4
|
-
|
|
5
|
-
Work like a senior engineer who owns this result, not a commentator on it.
|
|
6
|
-
- Evidence over assertion — back every claim with a file:line, or mark it an explicit assumption. Never state the unverified as fact.
|
|
7
|
-
- Read before you reason — read each required input end to end; when you lack basis, write "insufficient evidence" instead of filling the gap plausibly.
|
|
8
|
-
- Shortest sound path — chase the most likely cause first; don't re-verify what is settled or pad with restatement.
|
|
9
|
-
- Decide, don't survey — when options exist, give the trade-off and one recommendation, not an exhaustive list.
|
|
10
|
-
- Fit what's here — match the surrounding code and prose; size the response to the request.
|
|
11
|
-
- Own the synthesis — weigh worker outputs on evidence, not consensus; a better-grounded dissent outranks the majority.
|
|
12
|
-
|
|
13
3
|
## Overview
|
|
14
4
|
|
|
15
5
|
The lead orchestrates the selected AI workers against a prepared task bundle, collects their independent outputs, supervises convergence, and ensures the final report is produced. When `Report writer worker` is in the selected roster, that worker authors the final-report artifacts; the lead reviews and approves them. The lead never substitutes its own reasoning for a worker result and never bypasses a rostered report writer.
|
|
@@ -143,7 +133,7 @@ The sequence is fixed:
|
|
|
143
133
|
|
|
144
134
|
**The lead never invents a model.** Every role's model is read from `task-manifest.json` → `resultContract.requiredWorkerRoles[*].modelExecutionValue` (and the lead model metadata). A missing assignment is a manifest defect, not a license to fall back — see [team-contract](./team-contract.md) "Model Assignment Rules". The manifest is always populated at run-prep time by the CLI, which seeds these values from `OKSTRA_DEFAULT_*_MODEL` (`scripts/okstra_ctl/run.py`).
|
|
145
135
|
|
|
146
|
-
**Reading an assignment is not enough — the selected adapter must apply it at dispatch.** `dispatch_worker` receives the manifest assignment
|
|
136
|
+
**Reading an assignment is not enough — the selected adapter must apply it at dispatch.** `dispatch_worker` receives the complete manifest assignment. The selected runtime adapter passes `hostModelValue` to a `runner=native-session` host primitive or `modelExecutionValue` to a `runner=cli-wrapper` provider process without changing provider, role, or model. A missing or unsupported runner-specific mapping is a pre-dispatch contract failure, never a silent fallback.
|
|
147
137
|
|
|
148
138
|
The table below documents those prep-time seed values **for reference only** — it is NOT a lead-applied fallback:
|
|
149
139
|
|
|
@@ -152,19 +142,19 @@ The table below documents those prep-time seed values **for reference only** —
|
|
|
152
142
|
| Lead role | opus | -- | runtime-specific role label; orchestration + convergence supervision + final-report review/approval |
|
|
153
143
|
| Report writer worker | sonnet | report-writer-worker | `agents/workers/report-writer-worker.md` |
|
|
154
144
|
| Claude worker | opus | claude-worker | `agents/workers/claude-worker.md` |
|
|
155
|
-
| Codex worker | gpt-5.6-sol | codex-worker |
|
|
156
|
-
| Antigravity worker | gemini-3.1-pro | antigravity-worker |
|
|
145
|
+
| Codex worker | gpt-5.6-sol | codex-worker | duty + task instructions composed per invocation; deterministic `worker-dispatch` execution |
|
|
146
|
+
| Antigravity worker | gemini-3.1-pro | antigravity-worker | duty + task instructions composed per invocation; deterministic `worker-dispatch` execution |
|
|
157
147
|
|
|
158
|
-
|
|
148
|
+
Each analysis assignment follows its recorded `runner`. `runner=native-session` uses the host's native subagent primitive after `host-native-spec-link-gate`; `runner=cli-wrapper` uses the deterministic `okstra worker-dispatch` process boundary after `core-pre-dispatch` verification. No LLM transport wrapper sits in front of a provider CLI.
|
|
159
149
|
|
|
160
150
|
### Implementation phase: Executor binding
|
|
161
151
|
|
|
162
152
|
For `--task-type implementation` runs, the task bundle additionally pins one of `claude` / `codex` / `antigravity` as the Executor — the only worker permitted to mutate project files in that run. The binding is exposed in two canonical places:
|
|
163
153
|
|
|
164
|
-
- `instruction-set/analysis-profile.md` — top "Executor binding" block (provider,
|
|
165
|
-
- `runs/implementation/manifests/run-manifest-*.json` — `teamContract.executor` object (same
|
|
154
|
+
- `instruction-set/analysis-profile.md` — top "Executor binding" block (provider, display name, model, runner, and dispatch mode)
|
|
155
|
+
- `runs/implementation/manifests/run-manifest-*.json` — `teamContract.executor` object (the same binding plus `appliesTo: "implementation"`)
|
|
166
156
|
|
|
167
|
-
Lead MUST dispatch Edit/Write-bearing work only through the `
|
|
157
|
+
Lead MUST dispatch Edit/Write-bearing work only through that executor binding: use the host primitive with `hostModelValue` for `runner=native-session`, or `okstra worker-dispatch` with `modelExecutionValue` for `runner=cli-wrapper`. The other two providers still run as read-only verifiers in the same run; the executor's own provider is *also* dispatched separately as a verifier in a fresh session, so the diff is reviewed context-isolated. Session isolation is the primary self-review safeguard — same-model executor and same-provider verifier is acceptable in distinct sessions. A different model variant (e.g. executor=opus / Claude verifier=sonnet) is recommended but not mandatory.
|
|
168
158
|
|
|
169
159
|
Executor is chosen at run-prep time via `--executor <claude|codex|antigravity>` (or `OKSTRA_DEFAULT_EXECUTOR`, fallback `claude`); the model used by the executor is taken from the corresponding worker model flag (`--claude-model` / `--codex-model` / `--antigravity-model`). For CLI-backed executors, the underlying file mutation happens inside the executor CLI's own auto-edit mode (e.g. `codex exec --sandbox workspace-write`), not through the lead runtime's `write_artifact` operation.
|
|
170
160
|
|
|
@@ -267,7 +257,7 @@ The launch prompt's `## Run Logs (error-log wiring)` section gives Lead the reso
|
|
|
267
257
|
|
|
268
258
|
Workers are contractually required to extract these two lines and abort with `<WORKER>_ERRORS_PATH_MISSING` if either is absent (see each worker definition's "Path extraction (BLOCKING)" block). Omitting these headers reproduces the historical bug where every run's `errors-<task-type>-<seq>.jsonl` stayed empty (workers had only template placeholders).
|
|
269
259
|
|
|
270
|
-
After each worker terminates, BEFORE classifying its terminal status, verify the canonical result file exists at the absolute path resolved from the `**Result Path:**` header. If it is absent — or the
|
|
260
|
+
After each worker terminates, BEFORE classifying its terminal status, verify the canonical result file exists at the absolute path resolved from the `**Result Path:**` header. If it is absent — or the deterministic provider process returned `CODEX_RESULT_MISSING` / `ANTIGRAVITY_RESULT_MISSING` — re-dispatch the SAME worker once with the byte-identical prompt. Only after the second attempt also misses may the role be classified `error` with `--message "result-missing after 1 retry"`. Full rules: [team-contract](./team-contract.md) "Lead Redispatch Policy on Result-Missing".
|
|
271
261
|
|
|
272
262
|
After each worker terminates (any terminal status), if its errors sidecar exists, dump it to the run error log using the same resolved paths from the launch prompt:
|
|
273
263
|
|
|
@@ -280,13 +270,13 @@ okstra error-log append-from-worker \
|
|
|
280
270
|
|
|
281
271
|
`--agent`, `--agent-role`, and `--error-type` are **closed enums**, not free-form labels — the role names used elsewhere in these contracts (`Codex worker`, `Claude worker`) are rejected. Use exactly:
|
|
282
272
|
|
|
283
|
-
- `--agent` — `claude-worker` | `codex-worker` | `antigravity-worker` | `report-writer`
|
|
273
|
+
- `--agent` — `claude-worker` | `codex-worker` | `antigravity-worker` | `grok-worker` | `kimi-worker` | `report-writer`
|
|
284
274
|
- `--agent-role` — `lead` | `worker` | `report-writer`
|
|
285
275
|
- `--error-type` — `cli-failure` | `contract-violation` | `tool-failure`
|
|
286
276
|
|
|
287
277
|
For a lead-attributed event there is no value in the list above — the selected adapter names the lead identity to pass.
|
|
288
278
|
|
|
289
|
-
For Codex/Antigravity
|
|
279
|
+
For deterministic Codex/Antigravity provider processes: if the CLI returns non-zero, times out, or hits a rate limit, immediately call `append-observed` with the captured exit code, duration, message, and stderr excerpt. `append-observed` additionally requires `--phase`, `--command`, and `--command-kind`, so copy this form rather than trimming the one above:
|
|
290
280
|
|
|
291
281
|
```bash
|
|
292
282
|
okstra error-log append-observed \
|
|
@@ -303,7 +293,7 @@ okstra error-log append-observed \
|
|
|
303
293
|
|
|
304
294
|
Keep `--message` to the error actually observed — asserting that a sandbox or permission boundary blocked the call requires `--context-json` carrying `cause` plus both `causeEvidence` probes, and an unevidenced block claim in `--message` is rejected. If an `append-from-worker` dump is rejected for that reason, correct the offending sidecar entry and re-run the dump instead of skipping it: the dump aborts at the rejected entry, so every later entry in that sidecar never reaches the run log.
|
|
305
295
|
|
|
306
|
-
The
|
|
296
|
+
The deterministic dispatcher records this through its selected adapter — Lead does NOT need to re-record. Token usage is not inferred from dispatch return values; call `collect_usage` at the start of Phase 7.
|
|
307
297
|
|
|
308
298
|
## Phase 5.5: Convergence loop
|
|
309
299
|
|
|
@@ -181,6 +181,21 @@ Plan-body verification stays **lightweight** even under this posture — the `ve
|
|
|
181
181
|
|
|
182
182
|
## Round protocol (single round at default `maxRounds=1`)
|
|
183
183
|
|
|
184
|
+
Before each verifier call, write one task-instructions file under the current
|
|
185
|
+
run's `state/` directory and run `okstra agent-prompt materialize` with
|
|
186
|
+
`--audience reverification-worker`,
|
|
187
|
+
`--assignment-ref reverify/<workerId>`, the exact `--worker-id`, and
|
|
188
|
+
`--dispatch-kind reverify-r<N>`. Run `okstra agent-prompt verify` against the
|
|
189
|
+
returned `metadataPath` before dispatch and use the returned `promptPath`
|
|
190
|
+
without modification. Native-session calls use only `hostModelValue`; before
|
|
191
|
+
the host primitive, run `okstra agent-prompt record-dispatch` with the project
|
|
192
|
+
root, run manifest, metadata path, and `--enforcement-mode
|
|
193
|
+
host-native-spec-link-gate`, then run `okstra agent-prompt link-result` with
|
|
194
|
+
`--dispatch-id <invocationId>:attempt-1` and the result path before parsing it;
|
|
195
|
+
CLI-wrapper calls go through `okstra worker-dispatch` and consume only
|
|
196
|
+
`modelExecutionValue`. A missing or invalid invocation contract blocks the
|
|
197
|
+
round before any host or provider process starts.
|
|
198
|
+
|
|
184
199
|
1. Lead runs `okstra plan-items extract --data <data.json> --output <state>/plan-items-....json`, places the persisted `items[]` verbatim in every verifier prompt with the compact `subject` and lossless `payload`, then runs `okstra plan-items validate --data <data.json> --items <state>/plan-items-....json`. Dispatch only after that exact-match validation succeeds.
|
|
185
200
|
2. For each analyser worker in the roster (`claude`, `codex`, and `antigravity` if opted in), lead constructs a reverify prompt using the template in §"Plan-body reverify prompt" below.
|
|
186
201
|
3. Dispatch uses the same wrapper infrastructure as finding convergence, so the `--role-slug` is the same canonical `<role>-worker` that convergence uses — not a round-specific slug. Result file path: `runs/<task-type>/worker-results/<role>-worker-plan-verify-r<N>-implementation-planning-<seq>.md` (e.g. `codex-worker-plan-verify-r1-implementation-planning-003.md`). The `-worker-` token is load-bearing twice over: §"Plan-body reverify prompt" requires the same anchor headers as convergence, whose `**Audit sidecar path:**` is derived by `okstra_ctl.worker_artifact_paths.audit_sidecar_rel()` inserting `-audit-` after that token — a slug without it makes the header underivable and the helper raises. Record each `planItems[].verdicts[].worker` as the same `<role>-worker` string, because provenance compares it to this filename's prefix. **Enforced:** `tests/contract/test_reverify_dispatch_anchors.py` derives the sidecar from the documented name and re-extracts the prefix the provenance resolver uses.
|
|
@@ -357,7 +372,7 @@ The [convergence](./convergence.md) §"Required reverify output contract"
|
|
|
357
372
|
applies unchanged: append it verbatim after the response format below.
|
|
358
373
|
|
|
359
374
|
````
|
|
360
|
-
|
|
375
|
+
Perform plan-body verification for <task-key> (round 1).
|
|
361
376
|
|
|
362
377
|
## Instructions
|
|
363
378
|
|
|
@@ -23,12 +23,13 @@ Two `frontmatter` approval fields are always emitted with their unset default
|
|
|
23
23
|
## Phase 6 dispatch template (Report writer worker)
|
|
24
24
|
|
|
25
25
|
1. Resolve the Report writer worker assignment and all required prompt/result/error paths from the manifests.
|
|
26
|
-
2.
|
|
27
|
-
3.
|
|
28
|
-
4.
|
|
29
|
-
5.
|
|
26
|
+
2. Write a call-specific task-instructions file containing the anchor headers and audience-specific reading list.
|
|
27
|
+
3. Run `okstra agent-prompt materialize --audience report-writer --assignment-ref initial/report-writer --worker-id report-writer --dispatch-kind report-writer ...`, then run `okstra agent-prompt verify` against the returned `metadataPath`. Use the returned `promptPath` without appending role prose. A correction redispatch repeats this step with a fresh invocation ID and the same audience and assignment reference.
|
|
28
|
+
4. Emit the Phase 6 checkpoint.
|
|
29
|
+
5. For `runner=native-session`, first run `okstra agent-prompt record-dispatch` with the project root, run manifest, metadata path, and `--enforcement-mode host-native-spec-link-gate`, then call the host primitive with only the returned `hostModelValue`. After its result exists, run `okstra agent-prompt link-result` with `--dispatch-id <invocationId>:attempt-1` and the result path before accepting it. For `runner=cli-wrapper`, call `okstra worker-dispatch --workers report-writer`, which consumes `modelExecutionValue` and verifies the metadata before starting the provider process. Never combine this Phase 6 call with analysis workers.
|
|
30
|
+
6. Call `await_workers([handle])` and verify the data.json Result Path, rendered Markdown sibling, and worker-result pointer at Worker Result Path. Verify the separate heartbeat audit sidecar before accepting the run. **Enforced:** both dispatch adapters keep the three completion paths in `WorkerJob.completion_paths`, and `validators/validate_session_conformance.py` validates the audit sidecar.
|
|
30
31
|
|
|
31
|
-
The assignment
|
|
32
|
+
The complete assignment supplies both runner-specific model values and the prompt header in item 9 below. A native host uses `hostModelValue`; a deterministic provider process uses `modelExecutionValue`; the recorded `**Model:**` header remains the canonical assignment label. Missing or unsupported model resolution is a pre-dispatch contract failure; the common contract does not choose a runtime fallback.
|
|
32
33
|
|
|
33
34
|
The prompt MUST include, in this order at the top:
|
|
34
35
|
|
|
@@ -90,6 +91,20 @@ For an implementation-planning run, the Report writer worker owns the Phase 6 de
|
|
|
90
91
|
2. **Only when it passes and Report Language is not `en`**, dispatch the translator worker, which writes `final-report-<task-type>-<seq>.i18n.<lang>.json`.
|
|
91
92
|
3. Then run `report-finalize`.
|
|
92
93
|
|
|
94
|
+
For step 2, write translator-only task instructions and run `okstra
|
|
95
|
+
agent-prompt materialize` with `--audience translator`, `--assignment-ref
|
|
96
|
+
translator`, `--worker-id translator`, and `--dispatch-kind translator`. Run
|
|
97
|
+
`okstra agent-prompt verify` on the returned `metadataPath` before dispatch and
|
|
98
|
+
use the returned `promptPath` unchanged. A native-session call uses only
|
|
99
|
+
`hostModelValue`; first run `okstra agent-prompt record-dispatch` with the run
|
|
100
|
+
manifest, metadata path, and `--enforcement-mode
|
|
101
|
+
host-native-spec-link-gate`, then run `okstra agent-prompt link-result` with
|
|
102
|
+
`--dispatch-id <invocationId>:attempt-1` and the translation result before
|
|
103
|
+
accepting it. A CLI-wrapper call uses `okstra worker-dispatch` and its
|
|
104
|
+
`modelExecutionValue`. The host-native record links the accepted result to a
|
|
105
|
+
verified call specification but does not assert that Okstra observed the host's
|
|
106
|
+
actual prompt delivery.
|
|
107
|
+
|
|
93
108
|
**Never dispatch the translator before step 1.** The data.json is the English SSOT; a report-writer that authored it in the reader's language produces a translation *from that language into itself* — a full-cost, entirely useless artifact, and the run still fails at `check-source` afterwards. **Enforced:** `okstra report-translate extract` refuses to build a work list from a data.json over the Korean-prose limit, so a mis-ordered dispatch fails at the translator's first command instead of after it. When it does fail, the fix is a report-writer rewrite in English — discard the sidecar and `translation-source.json` produced from the Korean draft rather than editing them, because their English column is not English.
|
|
94
109
|
|
|
95
110
|
Phase 7 post-processing is then **one command**. `okstra report-finalize` owns the ordered sequence — it is the same code path the Codex lead adapter runs automatically, so a Claude-led run and a Codex-led run finalize identically:
|
|
@@ -19,8 +19,8 @@ Okstra tasks use one lead plus the exact worker assignments selected in the prep
|
|
|
19
19
|
|------|------|------|---------------|------|
|
|
20
20
|
| Lead | orchestration + convergence supervision + final-report review/approval | runtime-specific | -- | Does not author the final report when `Report writer worker` is rostered |
|
|
21
21
|
| Claude worker | Answer every brief question across feasibility, requirement interpretation, hidden assumptions, and alternatives — with file:line evidence | broad reasoning depth, hidden assumptions, execution-risk surfacing | claude-worker | `agents/workers/claude-worker.md` |
|
|
22
|
-
| Codex worker | Same core responsibility as Claude worker — identical questions, identical sections 1–5 | implementation realism, code-path implications, edge cases, technical trade-offs | codex-worker |
|
|
23
|
-
| Antigravity worker | Same core responsibility as Claude worker — identical questions, identical sections 1–5 | requirement interpretation, consistency, safety, alternative viewpoints | antigravity-worker |
|
|
22
|
+
| Codex worker | Same core responsibility as Claude worker — identical questions, identical sections 1–5 | implementation realism, code-path implications, edge cases, technical trade-offs | codex-worker | final prompt composed from the invocation duty and task instructions; CLI execution uses `worker-dispatch` |
|
|
23
|
+
| Antigravity worker | Same core responsibility as Claude worker — identical questions, identical sections 1–5 | requirement interpretation, consistency, safety, alternative viewpoints | antigravity-worker | final prompt composed from the invocation duty and task instructions; CLI execution uses `worker-dispatch` |
|
|
24
24
|
| Report writer worker | **Authors** the final-report file in Phase 6. NOT an analysis worker. | — | report-writer-worker | `agents/workers/report-writer-worker.md`. Excluded from Phase 4/5 and convergence |
|
|
25
25
|
|
|
26
26
|
**Model assignment has no default.** The model for every role comes from `resultContract.requiredWorkerRoles[*].modelExecutionValue` in `task-manifest.json` (and lead model metadata). There is no per-role hard-coded fallback — see "Model Assignment Rules" below.
|
|
@@ -32,8 +32,8 @@ Disjoint initial scopes are invalid triangulation. Every selected analysis worke
|
|
|
32
32
|
### Model Assignment Rules
|
|
33
33
|
|
|
34
34
|
1. `resultContract.requiredWorkerRoles` in `task-manifest.json` (and the lead model metadata) is the canonical source. There is no role-level fallback — a missing assignment is a manifest defect, not a license to invent one.
|
|
35
|
-
2.
|
|
36
|
-
3. **Dispatch-time enforcement (BLOCKING).** The selected adapter receives
|
|
35
|
+
2. Select the execution value from `runner`: `native-session` passes only `hostModelValue` to the host primitive, while `cli-wrapper` passes `modelExecutionValue` to the provider process. Both values remain recorded in the invocation contract; neither may be substituted for the other.
|
|
36
|
+
3. **Dispatch-time enforcement (BLOCKING).** The selected adapter receives the complete assignment and must apply the runner-specific value above. The adapter must fail before dispatch if it cannot apply the exact assignment; it must not inherit the lead model, change provider, or choose a nearby alias silently.
|
|
37
37
|
|
|
38
38
|
### Dynamic Worker Role Determination
|
|
39
39
|
|
|
@@ -51,7 +51,7 @@ Only workers selected from `recommendedWorkers` in `task-manifest.json` and `res
|
|
|
51
51
|
0. **Adapter-owned dispatch (BLOCKING).** Every worker start, await, retry, and shutdown goes through the selected runtime adapter. Core state records the outcome but never guesses a host primitive.
|
|
52
52
|
1. The lead is responsible for orchestration, convergence supervision, and final-report review/approval. It never overrides worker analysis and never bypasses a rostered Report writer worker.
|
|
53
53
|
2. `Report writer worker` is NOT an analysis worker. It is excluded from Phase 4/5 (initial analysis) and Phase 5.5 (convergence re-verification). It is spawned only in Phase 6 and is the **author** of the final-report file at `runs/<task-type>/reports/final-report-<task-type>-<seq>.md`.
|
|
54
|
-
3. When `Report writer worker` is in the roster, Lead MUST dispatch it in Phase 6. The only legal lead-authored fallback is when a dispatch was attempted and recorded a terminal status of `error` / `timeout` / `not-run` with a concrete logged reason. Speculative reasons such as "session resume constraint" or "team is no longer alive" are NOT valid — `dispatch_worker` can start a fresh one-shot assignment through the selected adapter.
|
|
54
|
+
3. When `Report writer worker` is in the roster, Lead MUST dispatch it in Phase 6 as a separate invocation after convergence. Omit it from Phase 4/5 analysis selection and pass `--workers report-writer` for a CLI-backed Phase 6 call. The only legal lead-authored fallback is when a dispatch was attempted and recorded a terminal status of `error` / `timeout` / `not-run` with a concrete logged reason. Speculative reasons such as "session resume constraint" or "team is no longer alive" are NOT valid — `dispatch_worker` can start a fresh one-shot assignment through the selected adapter. **Enforced:** `dispatch_core._validate_report_writer_isolation()` rejects every mixed analysis/report plan before process creation, and the default roster selectors exclude `report-writer`.
|
|
55
55
|
4. The assigned model for each role is maintained based on `resultContract.requiredWorkerRoles` in task-manifest.json and the lead model metadata.
|
|
56
56
|
5. Required roles must not be replaced by unnamed generic parallel workers.
|
|
57
57
|
6. Before dispatching any required worker, persist the exact worker prompt to the assigned current-run prompt history path under `runs/<task-type>/prompts/`.
|
|
@@ -152,11 +152,11 @@ Branch on the exit code, not the JSON: without `--wait`, `0` = every probe healt
|
|
|
152
152
|
|
|
153
153
|
## Lead Redispatch Policy on Result-Missing
|
|
154
154
|
|
|
155
|
-
After each worker
|
|
155
|
+
After each worker attempt returns (regardless of role), Lead MUST verify the canonical result file exists at the absolute path resolved from the `**Result Path:**` anchor header (against `**Project Root:**`). The check is identical for host-native workers and deterministic CLI processes.
|
|
156
156
|
|
|
157
157
|
**Triggers (any of):**
|
|
158
158
|
|
|
159
|
-
- The
|
|
159
|
+
- The deterministic provider process returned an explicit `*_RESULT_MISSING` sentinel.
|
|
160
160
|
- The result file is absent at the resolved absolute path even though the worker returned without a `*_RESULT_MISSING` sentinel — for example, claude-worker returned its final assistant message but never persisted the artifact, or the wrapper exited 0 and the codex/antigravity sub-agent forwarded raw stdout despite the contract.
|
|
161
161
|
- The result file exists but cannot be parsed (frontmatter unreadable, sections 1–5 entirely missing). A truncated file in the middle of section 5 is NOT covered here — it goes to the validator's regular `error` path, not the retry path.
|
|
162
162
|
- `okstra worker-liveness --team-state <path> --worker <id>` reports a **CLI-wrapper** worker (`codex` / `antigravity`) `did-not-launch` — neither `<prompt-path>.log` nor `<prompt-path>.status.json` exists after the persisted `startedAt` plus the launch grace (default 60s). The wrapper writes its status sidecar before invoking the CLI and hard-fails loudly with a distinct exit code on every argument check before that, so the absence of BOTH artifacts means the dispatch itself never reached the script. Without this trigger the only evidence was a lead noticing two missing files by eye, and the run paid the full polling cap for a worker that never started.
|
|
@@ -175,7 +175,7 @@ After each worker subagent returns (regardless of role), Lead MUST verify the ca
|
|
|
175
175
|
- Lead MUST log the deviation with the normal `contract-deviation` entry naming the removed or reordered reads and the measured evidence that motivated it (byte counts, elapsed time). An unlogged reading-plan change is a contract violation, not an exception.
|
|
176
176
|
- This exception never licenses changing what the worker is asked to *produce*. Narrowing the deliverable to make it finish is a contract violation.
|
|
177
177
|
|
|
178
|
-
**Logging.** Lead records the first attempt's `cli-failure` (already emitted by
|
|
178
|
+
**Logging.** Lead records the first attempt's `cli-failure` (already emitted by `worker-dispatch`) as-is. The retry, on success, is logged via the normal worker-completion path; on failure (second `*_RESULT_MISSING`), Lead records a single `contract-violation` entry with `--message "result-missing after 1 retry"` referencing both adapter dispatch-attempt ids and prompt-history paths.
|
|
179
179
|
|
|
180
180
|
**Diagnostic sidecar (advisory).** Every CLI-worker dispatch writes a heartbeat sidecar at `<prompt-path>.status.json` recording `started_ts`, `ended_ts`, `exit_code`, `duration_ms`, and the canonical `log_path` (written by `scripts/okstra_ctl/worker_runner.py`, which every provider entrypoint shares). Lead MAY read this sidecar when deciding whether the first attempt actually launched the CLI (stage=`exited`, `exit_code=0`, non-zero `duration_ms`) versus failed before reaching it (sidecar absent, or stage=`started` with no exit fields). A run that ended abnormally after launch — an error, a Ctrl-C, or the SIGTERM/SIGHUP a pane kill or session teardown sends — closes as stage=`exited` with a `failure` string and **no** `exit_code`; read that as a failed attempt, not as a success. A SIGKILL cannot be closed by anything, so a sidecar still reading stage=`started` is not evidence that the worker is alive. The sidecar is best-effort — its absence is NOT by itself a reason to skip the retry; the canonical trigger remains the missing result file.
|
|
181
181
|
|
|
@@ -240,14 +240,14 @@ wiring)` section (resolved by the okstra runtime via `paths.py`). If Lead
|
|
|
240
240
|
omits either header, the worker MUST return `<WORKER>_ERRORS_PATH_MISSING`
|
|
241
241
|
without proceeding.
|
|
242
242
|
|
|
243
|
-
- `cli-failure` events are recorded by
|
|
244
|
-
- **
|
|
245
|
-
- **Background dispatch + polling contract (
|
|
246
|
-
- Successful completion: return the
|
|
243
|
+
- `cli-failure` events are recorded by `worker-dispatch` directly to the run-level error log via `okstra error-log append-observed --error-type cli-failure ...` — NOT via the sidecar. The sidecar is a worker tool-failure channel only.
|
|
244
|
+
- **Provider-process invocation arity.** Every `okstra-<provider>-exec.sh` entrypoint takes the same three required positional arguments plus three optional ones: `<project-root> <model-execution-value> <prompt-path> [worktree-path] [role] [idle-timeout-seconds]`, optionally followed by `--presentation live|quiet`. `worker-dispatch` alone constructs this invocation from the verified `WorkerJob`; leads and host adapters do not assemble it. The fourth argument is mandatory for implementation, the fifth names the functional role, and the sixth controls the shared idle budget (1500s for executor/verifier, 600s otherwise). `live` is reserved for a pane backend; deterministic dispatch uses `quiet`.
|
|
245
|
+
- **Background dispatch + polling contract (CLI processes).** `worker-dispatch` starts the selected provider CLI through the adapter's asynchronous execution mapping and awaits the same handle until it reports terminal completion, capped at 30 minutes (1800s) of wall-clock elapsed time. The adapter's await operation is the wait primitive; do not add a standalone sleep or build shorter-sleep loops to bypass a host constraint. This rule applies in **every phase**. Recording responsibilities:
|
|
246
|
+
- Successful completion: return the provider process's accumulated stdout from the terminal await result. No log entry.
|
|
247
247
|
- Non-zero `exit_code`: record a `cli-failure` to the run-level error log with the real `exit_code` and observed `duration-ms`.
|
|
248
248
|
- Polling cap reached: perform a one-shot **mtime-grace check** on the wrapper's live log (`<prompt>.log`). If the log was written within the last 90 seconds and grace has not yet been applied, extend the cap from 1800s to 2100s and continue awaiting. Otherwise call the selected adapter's termination mapping, record `cli-failure` with `--exit-code 124 --duration-ms <observed_ms> --message "<wrapper> exceeded polling cap (grace=<applied|not-applied>, last_mtime_age=<n>s)"`, then return the language-specific `*_CLI_TIMEOUT` sentinel.
|
|
249
249
|
- The selected adapter owns runtime-session accounting for the full wrapper window; core retains only the observed start/end event boundaries.
|
|
250
|
-
- **No external timeout
|
|
250
|
+
- **No external timeout around `worker-dispatch`.** The deterministic dispatch owns BOTH timeout mechanisms: (1) the process polling cap (30min + optional 5min mtime grace), and (2) the shared runner's stream-idle watchdog. If the CLI produces nothing for `<idle-timeout-seconds>`, the runner terminates the process group and marks the status sidecar timed out. Lead MUST NOT layer an earlier host timeout around it.
|
|
251
251
|
- `contract-violation` events (C) are recorded by Lead via `okstra error-log append-observed --error-type contract-violation ...` after inspecting worker outputs.
|
|
252
252
|
- Lead's responsibility regarding the sidecar is to dump it to the run-level error log via `okstra error-log append-from-worker` after each worker terminates; Lead does not write into the sidecar.
|
|
253
253
|
|
|
@@ -8,7 +8,7 @@ exploration"):
|
|
|
8
8
|
file's body into the persisted executor prompt at dispatch time.
|
|
9
9
|
The `Coding-conventions preflight` heading below is the literal string the CLI
|
|
10
10
|
wrapper's "Executor preflight forwarding check" greps for in the persisted
|
|
11
|
-
prompt (
|
|
11
|
+
prompt through `prepare_agent_invocation()` before `worker-dispatch`.
|
|
12
12
|
-->
|
|
13
13
|
|
|
14
14
|
# Coding-conventions preflight (BLOCKING — runs before the first `Edit` / `Write`, and binds the TDD loop)
|
|
@@ -14,7 +14,7 @@ Same delivery paths as the other executor gates (see _implementation-executor.md
|
|
|
14
14
|
dispatch time (see `okstra_ctl.initial_prompt_materialization.materialize_initial_prompts()`).
|
|
15
15
|
The `Pre-commit diff review sweep` heading below is the literal string the CLI
|
|
16
16
|
wrapper's "Executor post-write gate forwarding check" greps for in the persisted
|
|
17
|
-
prompt (
|
|
17
|
+
prompt through `prepare_agent_invocation()` before `worker-dispatch`.
|
|
18
18
|
-->
|
|
19
19
|
|
|
20
20
|
# Pre-commit diff review sweep (BLOCKING — before the executor's final commit)
|
|
@@ -34,7 +34,7 @@ reaches it. Enforcement: the CLI wrapper refuses an Executor dispatch whose
|
|
|
34
34
|
persisted prompt lacks the heading `Coding-conventions preflight`
|
|
35
35
|
(`<SENTINEL_PREFIX>_PREFLIGHT_MISSING`) or either post-write heading
|
|
36
36
|
(`<SENTINEL_PREFIX>_POSTWRITE_GATE_MISSING`) — see
|
|
37
|
-
`
|
|
37
|
+
`prepare_agent_invocation()` before `worker-dispatch`.
|
|
38
38
|
-->
|
|
39
39
|
- **Stage discipline (when a preceding stage is `done`):** its code is behavior-frozen — you may call, extend, or compose with it, never change what it already does. The rule body travels with this prompt the same way the gates do (`prompts/profiles/_stage-discipline.md`); only its `implementation` bullet binds you, the `implementation-planning` one binds the planner. Declaration-level — no wrapper sentinel.
|
|
40
40
|
- **Non-interactive auto-execution (BLOCKING for `runner=cli-wrapper`).** A CLI-wrapper executor runs head-less — there is no human at the keyboard. Skills loaded during the run (tdd, coding-preflight, and others) contain "get user approval", "state your plan to the user and wait", or "ask before proceeding" gates written for interactive sessions; in this run those gates are **already satisfied** by the upstream `implementation-planning` approval (the plan this stage executes was human-approved). The executor MUST NOT stop to request approval, MUST NOT end its turn after only producing a plan, and MUST carry the stage through end-to-end — RED → GREEN → refactor → per-cycle commit → `### Stage Carry Evidence`. The ONLY skill step to skip is the interactive user-approval prompt itself; every other skill rule (TDD discipline, conventions, real-IO isolation) still binds. Stopping early for approval in a head-less run is the observed empty-exit failure (exit 0, no diff): treat it as `contract-violated`.
|
|
@@ -11,9 +11,11 @@
|
|
|
11
11
|
- antigravity — when added to the roster it joins the verifier set; when omitted only the default Claude+Codex verifiers participate. `--executor antigravity` requires `antigravity` in the roster: the direct CLI demands it explicitly in `--workers`, while the wizard adds it automatically when you pick antigravity as the executor.
|
|
12
12
|
- **Executor binding (resolved at run-prep time, fixed for this run):**
|
|
13
13
|
- Executor display name: `{{EXECUTOR_DISPLAY_NAME}}`
|
|
14
|
+
- Executor worker ID: `{{EXECUTOR_WORKER_ID}}`
|
|
14
15
|
- Executor provider: `{{EXECUTOR_PROVIDER}}` (validated against the provider registry's `executor` capability; chosen via `--executor` or `OKSTRA_DEFAULT_EXECUTOR`, default `claude`)
|
|
15
|
-
- Executor
|
|
16
|
-
- Executor
|
|
16
|
+
- Executor model: `{{EXECUTOR_MODEL_DISPLAY}}` (CLI launch value: `{{EXECUTOR_MODEL_EXECUTION_VALUE}}`; host-native launch value: `{{EXECUTOR_HOST_MODEL_VALUE}}`)
|
|
17
|
+
- Executor runner: `{{EXECUTOR_RUNNER}}`
|
|
18
|
+
- Executor dispatch mode: `{{EXECUTOR_DISPATCH_MODE}}`
|
|
17
19
|
- Wherever this profile mentions the `Executor`, it refers to the role bound above. **Every** analysis provider in the resolved roster is also dispatched as a verifier — including the executor's own provider, which runs *separately* as a fresh session with no shared context so no verdict comes from the session that wrote the diff (`_implementation-verifier.md` owns this rule). Verifier dispatches remain strictly read-only.
|
|
18
20
|
{{INCLUDE:_common-contract.md}}
|
|
19
21
|
{{INCLUDE:_stage-discipline.md}}
|
|
@@ -14,6 +14,7 @@ from okstra_ctl.adapters.hosts.capability_adapter import (
|
|
|
14
14
|
numbered_interaction_port,
|
|
15
15
|
)
|
|
16
16
|
from okstra_ctl.domain.host import HostDescriptor
|
|
17
|
+
from okstra_ctl.ports.host_model import NativeExecutionValueHostModelBindingPort
|
|
17
18
|
from okstra_ctl.registry.provider_registry import ProviderRegistry
|
|
18
19
|
|
|
19
20
|
|
|
@@ -42,6 +43,7 @@ def create_adapter(
|
|
|
42
43
|
worker_dispatch_port=PENDING_HOST_PORT,
|
|
43
44
|
usage_accounting_port=CliArtifactUsageAccountingPort(),
|
|
44
45
|
provider_registry: ProviderRegistry | None = None,
|
|
46
|
+
host_model_port=None,
|
|
45
47
|
) -> CapabilityHostAdapter:
|
|
46
48
|
return CapabilityHostAdapter(
|
|
47
49
|
DESCRIPTOR,
|
|
@@ -57,4 +59,8 @@ def create_adapter(
|
|
|
57
59
|
supported_functions=INTERACTION_FUNCTIONS,
|
|
58
60
|
detector=no_automatic_claim,
|
|
59
61
|
provider_registry=provider_registry,
|
|
62
|
+
host_model_port=host_model_port or NativeExecutionValueHostModelBindingPort(
|
|
63
|
+
DESCRIPTOR.id,
|
|
64
|
+
DESCRIPTOR.native_provider_id,
|
|
65
|
+
),
|
|
60
66
|
)
|
|
@@ -77,9 +77,9 @@ Render every numbered item as its option label followed by its description verba
|
|
|
77
77
|
| `read_artifacts` | Read the manifest-provided paths through the current Antigravity host file interface. |
|
|
78
78
|
| `write_artifact` | Write only core-authorized `.okstra/` artifacts and preserve their schemas. |
|
|
79
79
|
| `prompt_user` | Ask through the current host text/question interface and stop at approval gates until an explicit answer arrives. |
|
|
80
|
-
| `dispatch_worker` | Dispatch
|
|
80
|
+
| `dispatch_worker` | Verify each materialized invocation first. Dispatch `runner=native-session` through the current host with the returned `promptPath` and `hostModelValue`. Dispatch `runner=cli-wrapper` through `okstra worker-dispatch`, which consumes `modelExecutionValue`. **Not in a cmux run:** when `terminalBackend` is `cmux-pane`, the cmux adapter overrides this row. |
|
|
81
81
|
| `await_workers` | Await native host workers through the host primitive and CLI workers through their status sidecars, then verify terminal state and Result Paths. |
|
|
82
|
-
| `redispatch_worker` |
|
|
82
|
+
| `redispatch_worker` | Materialize and verify a fresh invocation, then start a fresh native worker or deterministic `worker-dispatch` attempt according to the persisted runner. |
|
|
83
83
|
| `shutdown_workers` | Perform host or process cleanup only for resources owned by this run. |
|
|
84
84
|
| `record_lead_event` | Append the required structured event to the manifest-provided `leadEventsPath`; emit the matching user-facing `PROGRESS:` line. |
|
|
85
85
|
| `collect_usage` | Collect host- or artifact-backed usage through the existing Okstra token-usage path; do not substitute another runtime's session log. |
|
|
@@ -92,6 +92,7 @@ Render every numbered item as its option label followed by its description verba
|
|
|
92
92
|
- Do not infer the current host from an installed `agy` binary. The `antigravity` runtime must come from the active host skill or an explicit runtime flag.
|
|
93
93
|
- Unsupported workers or unavailable models fail before dispatch; do not change the provider, model, or runner silently.
|
|
94
94
|
- Reverify and critic retries use fresh attempts and persist the core-supplied `dispatchKind`.
|
|
95
|
+
- Native calls first run `okstra agent-prompt record-dispatch` with the project root, run manifest, verified metadata path, and `--enforcement-mode host-native-spec-link-gate`; after the Result Path exists, run `okstra agent-prompt link-result` with `--dispatch-id <invocationId>:attempt-1` and that path before accepting it. This links the accepted result to a verified specification but does not prove the host-delivered bytes. CLI calls are pre-verified and recorded by `worker-dispatch`.
|
|
95
96
|
- Report-writer completion requires both the data Result Path and the worker-results audit path.
|
|
96
97
|
|
|
97
98
|
## Completion, cleanup, and resume
|