okstra 0.175.1 → 0.176.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/docs/architecture/storage-model.md +2 -2
- package/docs/architecture.md +22 -19
- package/docs/cli.md +18 -14
- package/docs/for-ai/skills/okstra-inspect.md +2 -3
- package/docs/for-ai/skills/okstra-rollup.md +1 -0
- package/docs/project-structure-overview.md +57 -56
- package/docs/task-process/README.md +11 -9
- package/docs/task-process/common-flow.md +13 -16
- package/docs/task-process/error-analysis.md +9 -10
- package/docs/task-process/final-verification.md +7 -7
- package/docs/task-process/implementation-planning.md +9 -9
- package/docs/task-process/implementation.md +6 -6
- package/docs/task-process/release-handoff.md +8 -7
- package/docs/task-process/requirements-discovery.md +8 -8
- package/package.json +1 -1
- package/runtime/BUILD.json +2 -2
- package/runtime/bin/lib/okstra/interactive.sh +12 -6
- package/runtime/bin/lib/okstra/usage.sh +3 -2
- package/runtime/bin/okstra-spawn-followups.py +4 -2
- package/runtime/prompts/launch.template.md +2 -2
- package/runtime/prompts/lead/context-loader.md +1 -2
- package/runtime/prompts/lead/okstra-lead-contract.md +3 -4
- package/runtime/prompts/lead/report-writer.md +15 -1
- package/runtime/prompts/lead/team-contract.md +16 -12
- package/runtime/prompts/profiles/_implementation-deliverable.md +1 -1
- package/runtime/prompts/profiles/_implementation-executor.md +0 -3
- package/runtime/prompts/profiles/_implementation-verifier.md +3 -7
- package/runtime/prompts/profiles/change-impact-analysis.md +9 -5
- package/runtime/prompts/profiles/error-analysis.md +14 -8
- package/runtime/prompts/profiles/feature-analysis.md +9 -5
- package/runtime/prompts/profiles/final-verification.md +10 -7
- package/runtime/prompts/profiles/implementation-option-selection.md +9 -5
- package/runtime/prompts/profiles/implementation-planning.md +13 -7
- package/runtime/prompts/profiles/implementation.md +9 -5
- package/runtime/prompts/profiles/improvement-discovery.md +10 -6
- package/runtime/prompts/profiles/project-analysis.md +9 -5
- package/runtime/prompts/profiles/requirements-discovery.md +15 -9
- package/runtime/prompts/wizard/prompts.ko.json +18 -1
- package/runtime/python/okstra_ctl/adapters/runtime/__init__.py +1 -0
- package/runtime/python/okstra_ctl/adapters/runtime/assembly.py +27 -0
- package/runtime/python/okstra_ctl/adapters/runtime/cli_wrapper.py +49 -0
- package/runtime/python/okstra_ctl/adapters/runtime/cmux.py +74 -0
- package/runtime/python/okstra_ctl/application/open_worker.py +29 -0
- package/runtime/python/okstra_ctl/assignment_resolver.py +27 -27
- package/runtime/python/okstra_ctl/dispatch_core.py +120 -100
- package/runtime/python/okstra_ctl/domain/host.py +3 -1
- package/runtime/python/okstra_ctl/domain/wizard/interaction.py +5 -0
- package/runtime/python/okstra_ctl/domain/worker_runtime.py +44 -0
- package/runtime/python/okstra_ctl/implementation_outcome.py +21 -5
- package/runtime/python/okstra_ctl/legacy_model_selection.py +115 -27
- package/runtime/python/okstra_ctl/manager_sync.py +4 -1
- package/runtime/python/okstra_ctl/next_phase.py +236 -0
- package/runtime/python/okstra_ctl/ports/worker_runtime.py +26 -0
- package/runtime/python/okstra_ctl/recap.py +4 -1
- package/runtime/python/okstra_ctl/render.py +46 -47
- package/runtime/python/okstra_ctl/role_requirements.py +28 -35
- package/runtime/python/okstra_ctl/rollup.py +4 -1
- package/runtime/python/okstra_ctl/run.py +8 -5
- package/runtime/python/okstra_ctl/stage_fix_carry.py +17 -1
- package/runtime/python/okstra_ctl/team.py +14 -18
- package/runtime/python/okstra_ctl/wizard.py +248 -56
- package/runtime/python/okstra_ctl/worker_prompt_body.py +18 -2
- package/runtime/python/okstra_ctl/worker_prompt_contract.py +40 -9
- package/runtime/python/okstra_ctl/worker_prompt_headers.py +16 -1
- package/runtime/python/okstra_ctl/workflow.py +18 -32
- package/runtime/python/okstra_ctl/worktree.py +3 -3
- package/runtime/python/okstra_ctl/worktree_registry.py +5 -4
- package/runtime/python/okstra_project/state.py +54 -6
- package/runtime/schemas/final-report-v2.0.schema.json +43 -5
- package/runtime/skills/okstra-inspect/facets/recap.md +2 -0
- package/runtime/skills/okstra-inspect/facets/report.md +1 -1
- package/runtime/skills/okstra-inspect/facets/status.md +15 -13
- package/runtime/skills/okstra-rollup/SKILL.md +1 -0
- package/runtime/skills/okstra-run/SKILL.md +5 -5
- package/runtime/templates/implementation-worker-preamble.md +3 -18
- package/runtime/templates/project-docs/task-index.template.md +0 -1
- package/runtime/templates/reports/html/macros/forms.html +5 -4
- package/runtime/templates/reports/html/tasks/final-verification.template.html +1 -1
- package/runtime/templates/reports/html/tasks/implementation.template.html +1 -1
- package/runtime/templates/worker-prompt-preamble.md +3 -36
- package/runtime/validators/validate-run.py +57 -99
|
@@ -24,16 +24,16 @@ flowchart TD
|
|
|
24
24
|
Type --> PlanPick[approved plan pick]
|
|
25
25
|
PlanPick --> Approved[approval marker confirm]
|
|
26
26
|
Approved --> Stage[stage pick<br/>whole-task or stage number]
|
|
27
|
-
Stage -->
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
27
|
+
Stage --> Leader[leader session read-only]
|
|
28
|
+
Leader --> RoleCount[role-count min..max<br/>omit uses recommended; skip if min==max]
|
|
29
|
+
RoleCount --> RoleModel[role-model provider/model per slot]
|
|
30
|
+
RoleModel --> RoleAdd[min=0 roles via role-add only<br/>default skip]
|
|
31
|
+
RoleAdd --> Extras[directive, related tasks, clarification]
|
|
32
|
+
Extras --> Confirm[confirmation]
|
|
33
33
|
Confirm --> Render[render-bundle]
|
|
34
34
|
```
|
|
35
35
|
|
|
36
|
-
|
|
36
|
+
Launch selection uses role slots and model refs only: leader is the current session (read-only), then each static role's count in `min..max` (default **recommended**; the count step is skipped when `min == max`), then one `provider/model` per slot. Roles with `min = 0` stay closed unless the user opens them with role-add (default skip). Duplicate model refs in the same role are rejected. There is no provider roster multi-pick and no `Use defaults / Customize` fork. Dynamic verifiers are not chosen at launch. `--workers` is a CLI compatibility input only, not a launch picker.
|
|
37
37
|
|
|
38
38
|
This phase does not ask for `base-ref` directly. The wizard selects whole-task or a single stage from the approved plan's Stage Map, and prepare resolves `VERIFICATION_TARGET` from the registry / `consumers.jsonl` / git state.
|
|
39
39
|
|
|
@@ -27,19 +27,17 @@ flowchart TD
|
|
|
27
27
|
Input -->|rerun| Prior[prior planning report via clarification-response]
|
|
28
28
|
Direction --> Worktree{active task worktree?}
|
|
29
29
|
Prior --> Worktree
|
|
30
|
-
Worktree -->|yes|
|
|
30
|
+
Worktree -->|yes| Leader[leader session read-only]
|
|
31
31
|
Worktree -->|no| BaseRef[base-ref pick/text]
|
|
32
|
-
BaseRef -->
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
Branch -->|yes| Models[role counts then role models; lead is a compatibility alias for leader]
|
|
37
|
-
Models --> Extras[directive, related tasks, clarification]
|
|
32
|
+
BaseRef --> Leader
|
|
33
|
+
Leader --> RoleCount[role-count min..max<br/>omit uses recommended]
|
|
34
|
+
RoleCount --> RoleModel[role-model provider/model per slot]
|
|
35
|
+
RoleModel --> Extras[directive, related tasks, clarification]
|
|
38
36
|
Extras --> Confirm
|
|
39
37
|
Confirm --> Render[render-bundle]
|
|
40
38
|
```
|
|
41
39
|
|
|
42
|
-
For a new plan, the wizard asks for a validated option-selection report and passes it as `--selected-direction`. A planning clarification rerun passes its own prior report through `--clarification-response`. The wizard currently does not ask about `--no-plan-verification`; on the okstra-run path, plan-body verification is prepared as enabled by default.
|
|
40
|
+
For a new plan, the wizard asks for a validated option-selection report and passes it as `--selected-direction`. A planning clarification rerun passes its own prior report through `--clarification-response`. Launch selection uses role slots and model refs only: planner count in `min..max` (default recommended), then one `provider/model` per slot. Duplicate model refs in the same role are rejected. There is no provider roster multi-pick. The wizard currently does not ask about `--no-plan-verification`; on the okstra-run path, plan-body verification is prepared as enabled by default.
|
|
43
41
|
|
|
44
42
|
## 3. prepare_task_bundle handling
|
|
45
43
|
|
|
@@ -59,12 +57,14 @@ sequenceDiagram
|
|
|
59
57
|
P->>P: provision/reuse task worktree
|
|
60
58
|
P->>R: _build_convergence_block()
|
|
61
59
|
R-->>M: convergence.planBodyVerification.enabled=true
|
|
62
|
-
P->>M: workflow nextRecommendedPhase
|
|
60
|
+
P->>M: workflow nextRecommendedPhase inherited, ready lowered to pending
|
|
63
61
|
P-->>W: prepared lead prompt
|
|
64
62
|
```
|
|
65
63
|
|
|
66
64
|
Prepare rejects a new plan without a selected-direction report. Comparison mode requires a valid `DIRECTION SELECTION` sidecar, while preselected-validation mode uses the confirmed upstream direction without one. The normalized snapshot binds the source report, source-data digest, option ID, direction body, requirements, and invariants.
|
|
67
65
|
|
|
66
|
+
Prepare does not name the next phase. It carries the inherited `workflow.nextRecommendedPhase` forward and lowers a `ready` pointer to `pending`, because this run has not finished and a `ready` pointer would read as an invitation to start the following phase. The pointer becomes `ready` at `implementation` only when Phase 7 projects it from this run's `implementationPlanning.outcome` of `plan-ready`.
|
|
67
|
+
|
|
68
68
|
## 4. lead execution flow
|
|
69
69
|
|
|
70
70
|
```mermaid
|
|
@@ -31,16 +31,16 @@ flowchart TD
|
|
|
31
31
|
Approved -->|no| Retry[re-prompt same step]
|
|
32
32
|
Approved -->|yes| Stage[stage multi-pick<br/>ready/active markers]
|
|
33
33
|
Stage --> Chain[render-args<br/>stage + chain-stages]
|
|
34
|
-
Chain -->
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
34
|
+
Chain --> Leader[leader session read-only]
|
|
35
|
+
Leader --> RoleCount[role-count min..max<br/>omit uses recommended; skip if min==max]
|
|
36
|
+
RoleCount --> RoleModel[role-model provider/model per slot]
|
|
37
|
+
RoleModel --> RoleAdd[min=0 roles via role-add only<br/>default skip]
|
|
38
|
+
RoleAdd --> Extras[directive, related tasks, clarification]
|
|
39
39
|
Extras --> Confirm
|
|
40
40
|
Confirm --> Render[render-bundle]
|
|
41
41
|
```
|
|
42
42
|
|
|
43
|
-
`
|
|
43
|
+
Launch selection uses role slots and model refs only: leader is the current session (read-only), then each static role's count in `min..max` (default **recommended**; the count step is skipped when `min == max`), then one `provider/model` per slot. Roles with `min = 0` stay closed unless the user opens them with role-add (default skip). Duplicate model refs in the same role are rejected. There is no provider roster multi-pick and no defaults-vs-customize fork. `executor` is only a compatibility alias for `implementer` in model refs; implementer slots are chosen through role-count / role-model. Dynamic verifiers are not chosen at launch. `--workers` is a CLI compatibility input only, not a launch picker. `stage_pick` is a multi-pick that shows done/in-progress/ready/waiting status. The selected stage set goes through dependency closure and topological sort into a `chain-stages` CSV, and each actual run executes only one of those stages.
|
|
44
44
|
|
|
45
45
|
The current okstra-run wizard path does not expose `--approve` that directly flips the approval checkbox. The plan file must already have a recognized approval marker.
|
|
46
46
|
|
|
@@ -27,19 +27,20 @@ flowchart TD
|
|
|
27
27
|
Type --> Plan[approved plan auto/pick]
|
|
28
28
|
Plan --> Scope[handoff stage pick<br/>whole-task or eligible stages]
|
|
29
29
|
Scope --> Worktree{active task worktree?}
|
|
30
|
-
Worktree -->|yes|
|
|
30
|
+
Worktree -->|yes| Leader[leader session read-only]
|
|
31
31
|
Worktree -->|no| BaseRef[base-ref pick/text]
|
|
32
|
-
BaseRef -->
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
32
|
+
BaseRef --> Leader
|
|
33
|
+
Leader --> RoleCount[role-count min..max<br/>omit uses recommended; skip if min==max]
|
|
34
|
+
RoleCount --> RoleModel[role-model provider/model per slot]
|
|
35
|
+
RoleModel --> RoleAdd[min=0 roles via role-add only<br/>default skip]
|
|
36
|
+
RoleAdd --> Extras[directive, related tasks, clarification]
|
|
36
37
|
Extras --> Template[PR template override?]
|
|
37
38
|
Template --> TemplateScope[save template to project/global?]
|
|
38
|
-
TemplateScope --> Confirm
|
|
39
|
+
TemplateScope --> Confirm[confirmation]
|
|
39
40
|
Confirm --> Render[render-bundle]
|
|
40
41
|
```
|
|
41
42
|
|
|
42
|
-
`release-handoff` has no worker roster
|
|
43
|
+
`release-handoff` has no analysis-worker dispatch. Launch selection still shows the leader session (read-only) and any applicable role-count / role-model steps; there is no provider roster multi-pick and no `Use defaults / Customize` fork. Dynamic verifiers are not chosen at launch. `--workers` is not a launch picker, and the runtime forces the worker list to empty. The wizard outcome's `renderArgs` includes `pr-template-path` only for release-handoff. Scope selection finishes before prepare, and the project/global save runs before `render-bundle` via the `config.set pr-template-path` action of `outcome.persistActions[]`. whole-task requires an accepted whole-task verification report, and for stage-group only the stages that were marked `verified` by an accepted single-stage verification in the Stage Lifecycle Snapshot but not yet covered by a `pr` become candidates.
|
|
43
44
|
|
|
44
45
|
Note that this phase is also a target of task worktree provisioning. The normal flow reuses the implementation/final-verification result of the same task-key. Starting a new task may create a new branch, and it is likely to be blocked at the entry gate's "implementation commit exists" condition.
|
|
45
46
|
|
|
@@ -31,18 +31,18 @@ flowchart TD
|
|
|
31
31
|
Keep -->|keep| Base
|
|
32
32
|
Keep -->|change/no brief| Brief
|
|
33
33
|
Type --> Base{active task worktree?}
|
|
34
|
-
Base -->|yes|
|
|
34
|
+
Base -->|yes| Leader[leader session read-only]
|
|
35
35
|
Base -->|no| BaseRef[base-ref pick/text]
|
|
36
|
-
BaseRef -->
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
36
|
+
BaseRef --> Leader
|
|
37
|
+
Leader --> RoleCount[role-count min..max<br/>omit uses recommended; skip if min==max]
|
|
38
|
+
RoleCount --> RoleModel[role-model provider/model per slot]
|
|
39
|
+
RoleModel --> RoleAdd[min=0 roles via role-add only<br/>default skip]
|
|
40
|
+
RoleAdd --> Extras[directive, related tasks, clarification]
|
|
41
|
+
Extras --> Confirm
|
|
42
42
|
Confirm --> Render[render-bundle]
|
|
43
43
|
```
|
|
44
44
|
|
|
45
|
-
|
|
45
|
+
Launch selection uses role slots and model refs only: leader is the current session (read-only), then each static role's count in `min..max` (default **recommended**; the count step is skipped when `min == max`), then one `provider/model` per slot. Roles with `min = 0` stay closed unless the user opens them with role-add (default skip). Duplicate model refs in the same role are rejected. There is no provider roster multi-pick and no `Use defaults / Customize` fork. Dynamic verifiers are not chosen at launch. `--workers` is a CLI compatibility input only, not a launch picker.
|
|
46
46
|
|
|
47
47
|
## 3. prepare_task_bundle handling
|
|
48
48
|
|
package/package.json
CHANGED
package/runtime/BUILD.json
CHANGED
|
@@ -192,8 +192,15 @@ PY
|
|
|
192
192
|
need_brief_val="$([[ -z "$BRIEF_PATH" ]] && printf '1' || printf '0')"
|
|
193
193
|
need_type_val="$([[ -z "$TASK_TYPE" ]] && printf '1' || printf '0')"
|
|
194
194
|
local autofill_output=""
|
|
195
|
-
|
|
195
|
+
# autofill 은 편의 기능이라 실패해도 run 을 끊으면 안 된다. heredoc 이
|
|
196
|
+
# okstra_ctl 을 import 하므로 설치가 깨지면 non-zero 로 끝나는데, okstra.sh 는
|
|
197
|
+
# `set -euo pipefail` 이라 이 평범한 대입 하나가 run 을 통째로 중단시킨다.
|
|
198
|
+
# 위 reconcile 호출(:189)이 같은 이유로 `|| true` 를 달고 있다.
|
|
199
|
+
autofill_output="$(PYTHONPATH="$OKSTRA_PYTHONPATH:${PYTHONPATH-}" python3 - "$manifest_path" "$need_brief_val" "$need_type_val" <<'PY'
|
|
196
200
|
import json, sys
|
|
201
|
+
|
|
202
|
+
from okstra_ctl import next_phase
|
|
203
|
+
|
|
197
204
|
path = sys.argv[1]
|
|
198
205
|
need_brief = sys.argv[2] == "1"
|
|
199
206
|
need_type = sys.argv[3] == "1"
|
|
@@ -208,15 +215,14 @@ task_type = ""
|
|
|
208
215
|
if need_brief:
|
|
209
216
|
brief = (data.get("taskBriefPath") or "").strip()
|
|
210
217
|
if need_type:
|
|
211
|
-
|
|
212
|
-
|
|
213
|
-
|
|
214
|
-
task_type = ""
|
|
218
|
+
# 포인터 판정은 Python SSOT 가 갖는다 — `ready` 가 아니면 빈 문자열이라
|
|
219
|
+
# 셸이 센티널 목록을 따로 들고 있을 이유가 없다.
|
|
220
|
+
task_type = next_phase.autofill_task_type(data)
|
|
215
221
|
|
|
216
222
|
print(f"BRIEF={brief}")
|
|
217
223
|
print(f"TYPE={task_type}")
|
|
218
224
|
PY
|
|
219
|
-
)"
|
|
225
|
+
)" || true
|
|
220
226
|
|
|
221
227
|
local manifest_brief=""
|
|
222
228
|
local manifest_type=""
|
|
@@ -74,8 +74,9 @@ optional arguments:
|
|
|
74
74
|
plans.
|
|
75
75
|
--task-key <project-id:task-group:task-id>
|
|
76
76
|
Shorthand for --project-id/--task-group/--task-id. When the matching task-manifest.json
|
|
77
|
-
exists, brief-path and task-type are auto-filled from it (taskBriefPath and
|
|
78
|
-
workflow.nextRecommendedPhase
|
|
77
|
+
exists, brief-path and task-type are auto-filled from it (taskBriefPath, and
|
|
78
|
+
workflow.nextRecommendedPhase.phase only while its status is \`ready\`).
|
|
79
|
+
Explicit flags always win.
|
|
79
80
|
|
|
80
81
|
options:
|
|
81
82
|
--render-only Render the host-neutral lead handoff prompt only. Do not launch a session.
|
|
@@ -50,6 +50,7 @@ from pathlib import Path
|
|
|
50
50
|
|
|
51
51
|
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
|
52
52
|
|
|
53
|
+
from okstra_ctl import next_phase # noqa: E402
|
|
53
54
|
from okstra_ctl.final_report_paths import final_report_markdown_path # noqa: E402
|
|
54
55
|
from okstra_ctl.paths import task_manifest_file # noqa: E402
|
|
55
56
|
from okstra_ctl.workflow import PHASE_SEQUENCE # noqa: E402
|
|
@@ -199,10 +200,11 @@ def _spawn_one(
|
|
|
199
200
|
"workflow": {
|
|
200
201
|
"currentPhase": suggested,
|
|
201
202
|
"currentPhaseState": "not-started",
|
|
202
|
-
"nextRecommendedPhase":
|
|
203
|
+
"nextRecommendedPhase": next_phase.make(
|
|
204
|
+
suggested, next_phase.STATUS_READY, "follow-up 으로 생성됨"
|
|
205
|
+
),
|
|
203
206
|
"phaseStates": {},
|
|
204
207
|
"awaitingApproval": False,
|
|
205
|
-
"routingStatus": "follow-up-spawned",
|
|
206
208
|
},
|
|
207
209
|
}
|
|
208
210
|
_write_manifest(task_manifest_file(task_root), manifest_payload)
|
|
@@ -18,11 +18,11 @@ For a new `implementation-planning` run, the plan-body sequence is initial verif
|
|
|
18
18
|
{{PHASE_ALLOWED_OUTPUTS}}
|
|
19
19
|
- Forbidden actions in this phase:
|
|
20
20
|
{{PHASE_FORBIDDEN_ACTIONS}}
|
|
21
|
-
- This run executes `{{WORKFLOW_CURRENT_PHASE}}` only. Do not start
|
|
21
|
+
- This run executes `{{WORKFLOW_CURRENT_PHASE}}` only. Do not start any later phase inside this run, even if the user says "proceed to the next step" or similar. Which phase comes next is not decided yet — this run's final report decides it.
|
|
22
22
|
{{STAGE_BATCH_DIRECTIVE}}
|
|
23
23
|
{{VERIFICATION_TARGET}}
|
|
24
24
|
{{STAGE_INTEGRATION}}
|
|
25
|
-
- Phase advancement requires a new okstra invocation launched with `--task-type
|
|
25
|
+
- Phase advancement requires a new okstra invocation, launched with an explicit `--task-type` after this run's final report is written and approved. The target of that run comes from the pointer this run's report authors into `workflow.nextRecommendedPhase`, and only when that pointer's `status` says it can be started. The lead must not write source code, run builds/migrations/deployments, or otherwise produce artifacts of a different phase from inside this run.
|
|
26
26
|
- See `Lifecycle Phase Boundaries` in the lifecycle core contract (`{{OKSTRA_LEAD_CONTRACT_PATH}}`) for the canonical rules and the phase-transition checklist.
|
|
27
27
|
|
|
28
28
|
{{TEAM_CREATION_GATE}}
|
|
@@ -46,9 +46,8 @@
|
|
|
46
46
|
| `workflow.currentPhaseState` | Current lifecycle phase state |
|
|
47
47
|
| `workflow.phaseStates` | Phase-by-phase lifecycle state map |
|
|
48
48
|
| `workflow.lastCompletedPhase` | Last completed lifecycle phase |
|
|
49
|
-
| `workflow.nextRecommendedPhase` | Next
|
|
49
|
+
| `workflow.nextRecommendedPhase` | Next-Phase Pointer — an object carrying the target phase, whether it can be started, and why. A projection of this report's Phase Routing; not authored here. Written per the Artifact Persistence Checklist in [report-writer](./report-writer.md); do not re-derive its rules here. |
|
|
50
50
|
| `workflow.awaitingApproval` | Approval wait marker |
|
|
51
|
-
| `workflow.routingStatus` | Routing decision status |
|
|
52
51
|
| `workflow.lastSafeCheckpoint` | Safe resume checkpoint metadata |
|
|
53
52
|
| `instructionSetPath` | Path to the `instruction-set/` **directory** containing `analysis-packet.md`, `analysis-profile.md`, `analysis-material.md`, `reference-expectations.md`, `task-brief.md`, `final-report-template.md`, and — only for task types whose host orchestration carries gates — `host-orchestration-rules.md` (see Step 4). Not a single-file path. |
|
|
54
53
|
| `referenceExpectationsPath` | config/deployment expectation artifact path |
|
|
@@ -47,8 +47,7 @@ Read-side inspection (`/okstra-inspect`) and scheduling (`/okstra-schedule-gen`)
|
|
|
47
47
|
- The `leader` owns orchestration, convergence supervision, and final-report review/approval. It does not author the final-report file when `Report writer worker` is in the roster. `lead` is a compatibility alias for `leader` and must not be written on new artifacts.
|
|
48
48
|
- Dispatch consumes stored role executions, not provider-named worker IDs. Canonical roles are `leader`, `analyser`, `critic`, `designer`, `planner`, `implementer`, `verifier`, `report-writer`, and `translator`. `executor` is a compatibility alias for `implementer`.
|
|
49
49
|
- Pane titles and operational rows use the stored `executionLabel`. Do not rebuild that label from a provider name or model string.
|
|
50
|
-
-
|
|
51
|
-
- `Report writer worker`, when in the roster, is the **author** of the final-report file. Lead reviews the draft and may request a revision via a follow-up dispatch, but MUST NOT write the report itself as a "shortcut". The only legal lead-authored fallback is when a Report writer worker dispatch was actually attempted and recorded a terminal status of `error`/`timeout`/`not-run` with an explicit reason in team-state — see [report-writer](./report-writer.md) "Lead-authored fallback".
|
|
50
|
+
- `report-writer`, when in the roster, is the **author** of the final-report file. Lead reviews the draft and may request a revision via a follow-up dispatch, but MUST NOT write the report itself as a "shortcut". The only legal lead-authored fallback is when a Report writer worker dispatch was actually attempted and recorded a terminal status of `error`/`timeout`/`not-run` with an explicit reason in team-state — see [report-writer](./report-writer.md) "Lead-authored fallback".
|
|
52
51
|
- "Session resume", "team is no longer alive", and similar are NOT valid reasons to skip Report writer worker dispatch — see [report-writer](./report-writer.md) "Resume-safe dispatch".
|
|
53
52
|
- A shell command the lead runs must not be able to ask a question. The lead's shell is the user's own, where `cp`, `mv`, and `rm` are commonly aliased to their `-i` form; the confirmation that alias raises has nobody to answer it, so the call hangs until it is killed — observed as a `cp` over an existing state file stalling a whole self-fix round. Invoke these as `command cp` / `command mv` / `command rm`, which skips alias expansion and leaves the tool's own behaviour untouched. `-f` is not a substitute: it changes what the tool does on failure (`rm -f` reports success on a path that never existed).
|
|
54
53
|
- If the brief is incomplete, continue with explicit uncertainty markers rather than fabricating confidence.
|
|
@@ -73,14 +72,14 @@ Phase-transition checklist (lead, end of run):
|
|
|
73
72
|
|
|
74
73
|
1. Confirm the current phase's required outputs are complete and recorded in the final report.
|
|
75
74
|
2. Set `workflow.phaseStates.<currentPhase>.state = "completed"` in `task-manifest.json` (validator does this when the run passes; verify the value).
|
|
76
|
-
3. Update `workflow.lastCompletedPhase
|
|
75
|
+
3. Record **Phase Routing** in the final report — it is the source the Next-Phase Pointer is projected from, and Phase 7 validation recomputes the pointer from it. Update `workflow.lastCompletedPhase`. The Next-Phase Pointer (`workflow.nextRecommendedPhase`) itself is written per the Artifact Persistence Checklist in [report-writer](./report-writer.md), which owns its shape and its rules.
|
|
77
76
|
4. **Do NOT start the next phase inside the current run.** A new okstra invocation with the new `--task-type` is the only legal way to advance.
|
|
78
77
|
|
|
79
78
|
User-utterance interpretation rule:
|
|
80
79
|
|
|
81
80
|
- "proceed to the next step" / "move on to the next step" / equivalent phrases are scoped to **the current phase only**. Interpret them as "produce the remaining outputs of the current phase," never as "start the next lifecycle phase."
|
|
82
81
|
- If the current phase's outputs are already complete and the user clearly wants to advance, reply with the phase-transition checklist above and the exact next-run command. Wait for explicit user confirmation before any action that belongs to the next phase.
|
|
83
|
-
- If `nextRecommendedPhase` is `implementation-planning`, the next run produces a **plan**, not code. The next run after that is `implementation`.
|
|
82
|
+
- If the Next-Phase Pointer target (`nextRecommendedPhase.phase`) is `implementation-planning`, the next run produces a **plan**, not code. The next run after that is `implementation`.
|
|
84
83
|
|
|
85
84
|
## Progress reporting (BLOCKING)
|
|
86
85
|
|
|
@@ -357,6 +357,8 @@ Every field MUST anchor its claim with at least one evidence reference — a `pa
|
|
|
357
357
|
2. **Evidence and Detailed Analysis** — primary evidence rows (file path, line, snippet); secondary evidence / alternate interpretations. If `reference-expectations.md` lists explicit expected values, record match/gap per row.
|
|
358
358
|
- **Error-analysis diagnosis and routing.** When `header.taskType` is `error-analysis`, populate the required `errorAnalysis` object. Copy `errorAnalysis.symptomVerbatim` byte-for-byte from the symptom stated in the brief's `Source Material`; do not paraphrase it. Every `causeCandidates[]` row includes the full `supportingEvidence`, `falsifyingEvidenceChecked`, `confidence`, and `disproveWith` fields. When a candidate is a step in a propagation chain rather than a competing explanation — the analysis calls it a downstream step, a second stage, or a consequence of another candidate — set its `downstreamOf` to the ids of the candidates immediately upstream of it; leave the field absent for a candidate that stands on its own. Every id listed MUST be another candidate in the same report, no row may name itself, and the links MUST NOT form a cycle; `validators/validate-run.py::_validate_cause_chain` rejects all three. This is the only place the chain is machine-readable — prose calling a candidate "the second step of the chain" while `downstreamOf` is absent leaves the report's figure claiming the candidates are alternatives. Route `errorAnalysis.routing.nextTaskType=implementation-option-selection` with `direction=begin-option-selection`, or route `errorAnalysis.routing.nextTaskType=error-analysis` with `direction=continue-investigation`; no other pairing is valid. `verdictCard.nextStep`, `finalVerdict.nextStep`, the first `recommendedNextSteps` action and command, and the unique `followUpTasks` row whose `origin` is `phase-continuation` MUST all point to the same `errorAnalysis.routing.nextTaskType` target. The schema enforces only the presence of a `phase-continuation` row; `validators/validate-run.py::_validate_error_analysis_consistency` enforces exact target agreement and uniqueness.
|
|
359
359
|
- **Implementation-option-selection comparison.** When `header.taskType` is `implementation-option-selection`, populate `implementationOptionSelection` from the converged direction-selection findings. Preserve every merged or rejected raw candidate in `candidateAudit`, and put at most three selectable candidates in `rankedOptions`. Each displayed candidate carries its requirement coverage, scope commitments, criterion scores, feasibility votes, safety blockers, unresolved feasibility facts, planning invariants, and exact coverage summary. In each displayed candidate, `expectedChangeAreas` names direction-level change surfaces, never exact file paths or an exact file list. `expectedVerification` names direction-level verification signals, never a stage list or executable test commands. `schemas/final-report-v2.0.schema.json` enforces the displayed-summary constants and the three-option cap; semantic recalculation belongs to `validators/validate-run.py`.
|
|
360
|
+
- **Routing.** `implementationOptionSelection.routing` is a required **string enum** — not an object — with exactly three values: `implementation-planning`, `pending-direction-selection`, `blocked`. It is the only field in this report that records where the task goes next, and Phase 7 projects `workflow.nextRecommendedPhase` from it (`scripts/okstra_ctl/next_phase.py`): `implementation-planning` becomes a `ready` pointer naming that phase, while `pending-direction-selection` and `blocked` become `pending` and `blocked` pointers carrying no phase. Only the first proposes a next run.
|
|
361
|
+
- **The value is determined by this run's mode and candidate set, not chosen freely.** `preselected-validation` mode routes to `implementation-planning` — the direction was already selected and this run only validated it. `candidate-comparison` mode that displays any candidate routes to `pending-direction-selection` — the user still owes the direction pick, so a comparison never routes straight to planning. `blocked` is legal only when no valid candidate exists at all, and is required in that case. **Enforced:** `scripts/okstra_ctl/implementation_options.py::validate_implementation_option_selection` rejects all three mismatches (`validated preselected direction must route to implementation-planning`, `candidate-comparison with options must await direction selection`, `routing must be blocked only when no valid options exist` / `routing may be blocked only when no valid options exist`).
|
|
360
362
|
- **Implementation-planning direction branch.** When `implementationPlanning.planningContract == "selected-direction"`, read `selectedDirectionRef` and the snapshot before authoring. Materialize the snapshot into `directionRealization`, stages, validation, rollback, and bidirectional original-requirement links. Author exactly one `P-Dir-1`; its payload is the complete `directionRealization`. Its verification covers the core mechanism, architecture boundaries, planning invariants, and any hidden direction change against `selectedDirectionRef`. Do not author Option Candidates, candidate scores, a Recommended Option, or user candidate-selection fields. When current evidence requires changing the direction, author `outcome: "direction-invalidated"` and omit the execution plan. Legacy candidate-comparison reruns retain `P-Opt-*`, Option Candidates, trade-off, and Recommended Option semantics.
|
|
361
363
|
|
|
362
364
|
```json
|
|
@@ -370,6 +372,12 @@ Every field MUST anchor its claim with at least one evidence reference — a `pa
|
|
|
370
372
|
```
|
|
371
373
|
|
|
372
374
|
- **Implementation-option-selection is non-terminal.** Its `followUpTasks` includes a `phase-continuation` row with `autoSpawn: "no"` and `priority: "P0"`; the schema's non-terminal conditional enforces row presence.
|
|
375
|
+
- **Implementation and final-verification routing.** Both phases record where the task goes next in a `routingRecommendation` **object** with exactly two fields: `target` is one enum value, `rationale` is the sentence that justifies it. Neither field takes free-form routing prose, and a target named only in the prose does not count — Phase 7 projects `workflow.nextRecommendedPhase` from `target` alone (`scripts/okstra_ctl/next_phase.py`), so the value you write there is the route the task actually takes.
|
|
376
|
+
- `implementation.routingRecommendation.target` is one of `final-verification`, `error-analysis`, `implementation-planning`, `implementation`. Pick `final-verification` when this stage's plan items landed and validation passed; `error-analysis` when a failure's cause is not understood; `implementation-planning` when the approved plan itself no longer fits the evidence; `implementation` when work remains inside this stage (the next run is a fix run).
|
|
377
|
+
- `finalVerification.routingRecommendation.target` is one of `release-handoff`, `release-handoff(stage-group)`, `error-analysis`, `implementation-option-selection`, `implementation-planning`, `implementation`, `done`. Both `release-handoff` values require the `accepted` verdict, and plain `release-handoff` additionally requires `verificationScope` `whole-task` — a `single-stage` accepted run routes to `release-handoff(stage-group)` instead. `done` ends the lifecycle here. `error-analysis` / `implementation-option-selection` / `implementation-planning` follow the cause-vs-direction-vs-plan split of the verdict token table above.
|
|
378
|
+
- `rationale` is one or two sentences on why that target and nothing else, citing the blocker ids or evidence rows behind the choice. It is the only free-form half of the field; the digest sections still point here for the full reasoning.
|
|
379
|
+
- `releaseHandoff.routingRecommendation` is unchanged — it stays a single prose field, because `release-handoff` is terminal and nothing projects a next phase from it.
|
|
380
|
+
- **Enforced:** `schemas/final-report-v2.0.schema.json` rejects a `target` outside the enum, a missing `rationale`, and any string value in either field; `validators/validate-run.py::_validate_final_verification_consistency` rejects a final-verification report whose `routingRecommendation.target` is absent, and rejects the verdict↔routing and scope↔routing combinations named above.
|
|
373
381
|
3. **Recommended Next Steps** — prioritized actions. After Phase 7's follow-up spawner runs, append a row per newly created task-key (see "Phase 6 → Phase 7 execution sequence" above). **Approval-gate consistency:** when §1 carries any `Blocks: approval` row with `Status` ∈ {open, answered}, the Verdict Card `Next Step` and the first recommended step MUST point to the clarification rerun (`resume-clarification` of the SAME task-type) — never to "flip frontmatter `approved: true` → jump straight to `implementation`". Run-prep enforces this gate (`run.py _validate_approved_plan` fail-closes on those rows and on a blocking data.json `gateResult`), so a direct-implementation next-step is an instruction the reader cannot actually follow. **Cross-project pointer rule:** for cross-project dependencies (another repo / a different top-level deployment module / a published package), `crossProjectDependencies` (§5.4 Cross-Project Dependencies) is authoritative — do NOT duplicate that substance (prerequisite work / verification signals / handoff) into `recommendedNextSteps`; put only a one-line pointer to that section (no double-recording).
|
|
374
382
|
4. **Follow-up Tasks** — auto-spawn-eligible table. Each row drives `okstra-spawn-followups.py`; see template §4 for the row schema.
|
|
375
383
|
5. **Missing Information and Risks** — uncertain / "I don't know" items. `implementation-planning` adds §5.5 (see heading contract below); `release-handoff` adds §5.6.
|
|
@@ -417,7 +425,13 @@ Persistence steps that must be performed in Phase 7:
|
|
|
417
425
|
- [ ] 4. **Update task-manifest.json**: Reflect task-level status and workflow lifecycle metadata
|
|
418
426
|
- Update `workCategory` if the run produced a confident classification
|
|
419
427
|
- Update `workflow.currentPhase`, `workflow.currentPhaseState`, `workflow.lastCompletedPhase`, and `workflow.phaseStates`
|
|
420
|
-
-
|
|
428
|
+
- Write `workflow.nextRecommendedPhase` as an object with exactly three string fields — `phase`, `status`, `rationale`. **This checklist item is the canonical statement of that field**; the lead contract, the context loader and every okstra skill point here instead of restating it, so a change to the rule belongs in this bullet.
|
|
429
|
+
- `status` is one of `ready` (the named phase can be started now), `pending` (this run did not settle where the task goes next), `blocked` (something outside this run must change before any phase can start), or `terminal` (the lifecycle ends here; there is no next phase).
|
|
430
|
+
- `phase` carries a lifecycle phase name only when `status` is `ready`; under the other three, write the empty string. This is an authoring rule for the value **you** write, not a constraint the struct enforces or a shape you can rely on when reading. `prepare` lowers a `ready` pointer to `pending` and keeps its `phase` (`scripts/okstra_ctl/render.py::_derive_next_recommended_phase`), so every in-flight task's manifest holds a non-`ready` pointer that still names a phase. Never infer launchability from a non-empty `phase` — read `status`, which is the one field that answers it.
|
|
431
|
+
- `rationale` is one sentence saying why. It is the only free-form field, and it is the only part of the pointer that survives a correction.
|
|
432
|
+
- **`phase` and `status` MUST agree with this report's own routing field.** Which field that is depends on the task-type: `requirementsDiscovery.routing.nextTaskType`, `errorAnalysis.routing.nextTaskType`, `implementationOptionSelection.routing`, `implementationPlanning.outcome`, `implementation.routingRecommendation.target`, or `finalVerification.routingRecommendation.target` (see "Implementation and final-verification routing" above for those two enums). Two task-type groups have no routing field that Phase 7 projects from: `release-handoff` (its `routingRecommendation` is prose that nothing reads) is always `terminal`, and the analysis sidetracks — `improvement-discovery`, `project-analysis`, `feature-analysis`, `change-impact-analysis` — are always `pending`, whatever a per-candidate recommendation inside the report says. One routing value is not a phase name: when `finalVerification.routingRecommendation.target` is `release-handoff(stage-group)`, write `phase` as `release-handoff`. The parenthesised part names the handoff's scope, not a different phase, and Phase 7 projects it that way — writing the parenthesised form here records a correction against a report that was right.
|
|
433
|
+
- **Enforced:** Phase 7 validation recomputes `phase` and `status` from that routing field (`scripts/okstra_ctl/next_phase.py::project`). On a passing run whose authored pair disagrees with the recomputed pair, `validators/validate-run.py` overwrites both with the recomputed values, replaces your `rationale` with a pointer sentence, and preserves what you wrote under `workflow.nextRecommendedPhaseCorrection.authored`. A run whose validation fails ends with the pointer `blocked`, keeping your `rationale`. So the routing field is what actually moves the task — a pointer authored against the report body changes nothing but the audit trail.
|
|
434
|
+
- Update `workflow.awaitingApproval`
|
|
421
435
|
- Update `workflow.lastSafeCheckpoint` to the best resume point for the current task
|
|
422
436
|
- [ ] 5. **Update task-index.md**: Refresh human-readable summary
|
|
423
437
|
- [ ] 6. **Generate final status file**: `runs/<task-type>/status/final-<task-type>-<seq>.status` (if necessary)
|
|
@@ -13,19 +13,23 @@ Okstra tasks use one lead plus the exact worker assignments selected in the prep
|
|
|
13
13
|
|
|
14
14
|
### Role Definitions
|
|
15
15
|
|
|
16
|
-
|
|
16
|
+
The start screen assigns a **model ref** (`provider/model`) to each **canonical role** slot: `leader`, `analyser`, `critic`, `designer`, `planner`, `implementer`, `verifier`, `report-writer`, `translator`. A provider name is the front of a model ref, not a role. Every `analyser` in the run shares an identical core responsibility. Specialization is additive — it lives in optional Section 6 of the worker output, NOT in differentiated core questions. Cross-verification only converges if every rostered analyser answers the same questions against the same brief.
|
|
17
17
|
|
|
18
|
-
|
|
|
19
|
-
|
|
20
|
-
|
|
|
21
|
-
|
|
|
22
|
-
|
|
|
23
|
-
|
|
|
24
|
-
|
|
|
18
|
+
| Canonical role | Core responsibility | Notes |
|
|
19
|
+
|------|------|------|
|
|
20
|
+
| `leader` | orchestration + convergence supervision + final-report review/approval | Does not author the final report when `report-writer` is rostered |
|
|
21
|
+
| `analyser` | Answer every brief question across feasibility, requirement interpretation, hidden assumptions, and alternatives — with file:line evidence | Same sections 1–5 for every model ref |
|
|
22
|
+
| `critic` | Audit coverage gaps or challenge acceptance, per the attached duty | Not an analysis voter |
|
|
23
|
+
| `designer` | Compare or validate implementation directions | Used by `implementation-option-selection` |
|
|
24
|
+
| `planner` | Produce an executable plan without writing the implementation | Used by `implementation-planning` |
|
|
25
|
+
| `implementer` | Sole change author for one approved stage | Used by `implementation` |
|
|
26
|
+
| `verifier` | Independent review of the assigned duty's subject | Duty varies by task type |
|
|
27
|
+
| `report-writer` | **Authors** the final-report file in Phase 6 | Excluded from Phase 4/5 and convergence |
|
|
28
|
+
| `translator` | Translates the designated sidecar only | Phase 7 |
|
|
25
29
|
|
|
26
|
-
**
|
|
30
|
+
**Dispatch does not invent a missing model assignment.** Launch-time empty slots are filled by the model default chain before the run starts. At dispatch the model for every role comes from `resultContract.requiredWorkerRoles[*].modelExecutionValue` in `task-manifest.json` (and lead model metadata). There is no per-role hard-coded fallback — see "Model Assignment Rules" below.
|
|
27
31
|
|
|
28
|
-
**Dispatch-prompt invariant.** Lead's dispatch prompt body for
|
|
32
|
+
**Dispatch-prompt invariant.** Lead's dispatch prompt body for every rostered analysis worker MUST be byte-identical except for the role label and any wrapper-specific path headers (e.g. `**Worktree:**`, `**Errors sidecar path:**`). The role label is the ONLY identity form the normalizer erases, and it erases exactly the label `worker_prompt_body.analysis_worker_label` renders for the run's own worker ids — so a provider outside that function's display map (`grok`, `kimi`, an installed adapter) is covered as it comes. Naming the worker any other way (a model name, a host name, a provider's product name) survives normalization and fails the equality group before publication. **Enforced:** `okstra_ctl.worker_prompt_contract.normalise_analysis_prompt`, pinned by `tests/contract/test_analysis_prompt_identity_normalization.py`. Lead MUST NOT bias the brief by inserting per-worker emphasis sentences ("you focus on X") into the body. Bias-by-prompt reproduces the historical failure mode where Claude commented only on assumptions, Codex only on code paths, and Antigravity only on requirements — leaving convergence with nothing to converge on.
|
|
29
33
|
|
|
30
34
|
Disjoint initial scopes are invalid triangulation. Every selected analysis worker owns the same common verification requirements; provider diversity supplies independent observations, not separate coverage slices. Do not shard the common scope by worker, provider, or model. Worker-specific depth belongs only in the non-voting Specialization Lens after the shared analysis is complete.
|
|
31
35
|
|
|
@@ -65,7 +69,7 @@ Only workers selected from `recommendedWorkers` in `task-manifest.json` and `res
|
|
|
65
69
|
|
|
66
70
|
`okstra_ctl.initial_prompt_materialization` is the canonical owner of roster-derived initial prompt rendering, validation, and immutable publication. Code-backed `dispatch_worker` mappings leave missing roster prompts to `materialize_initial_prompts()` and pass the selected adapter's declared `initialPromptDeliveryMode`; they do not compose or overwrite those prompts themselves. Native in-process dispatch uses the same headers and the adapter's declared `lazy-path-reference` mode, persists the prompt before dispatch, and remains outside reverify and critic handling.
|
|
67
71
|
|
|
68
|
-
Every worker prompt MUST start with the anchor headers rendered by `okstra_ctl.worker_prompt_headers.worker_prompt_headers()` (the generating SSOT — never hand-author or reorder them). Every persisted initial prompt also carries exactly one non-empty `**Prompt Delivery Mode:** <mode>` header whose value is `eager-include` or `lazy-path-reference`. Redispatch reuses that persisted initial prompt byte-for-byte; it never regenerates or overwrites it. The generated absolute `**Audit sidecar path:**` is derived by `audit_sidecar_rel()` from the canonical worker result path; workers write to that header and never synthesize a `runs/<task-type>/...` destination. Their meaning and extraction rules for workers are
|
|
72
|
+
Every worker prompt MUST start with the anchor headers rendered by `okstra_ctl.worker_prompt_headers.worker_prompt_headers()` (the generating SSOT — never hand-author or reorder them). Every persisted initial prompt also carries exactly one non-empty `**Prompt Delivery Mode:** <mode>` header whose value is `eager-include` or `lazy-path-reference`. Redispatch reuses that persisted initial prompt byte-for-byte; it never regenerates or overwrites it. The generated absolute `**Audit sidecar path:**` is derived by `audit_sidecar_rel()` from the canonical worker result path; workers write to that header and never synthesize a `runs/<task-type>/...` destination. Their meaning and extraction rules for workers are owned by `okstra_ctl.worker_prompt_headers.worker_prompt_headers()`. Phase-specific extra headers (implementation worktree, final-verification target snapshot, improvement-discovery grilling log) are emitted there too.
|
|
69
73
|
|
|
70
74
|
The `**Coding preflight pack:**` anchor is emitted only for the `implementation-executor` and `implementation-verifier` audiences. An implementation report-writer never receives it. A pane display role named `verifier` during `final-verification` is still an initial analysis worker; it neither receives implementation conventions nor becomes a Phase 5.5 reverify dispatch. Reverify is identified by the `-reverify-r<N>-` prompt/result path contract in [convergence](./convergence.md).
|
|
71
75
|
|
|
@@ -99,7 +103,7 @@ Audience-scoped file enumeration (performance optimization — mandatory):
|
|
|
99
103
|
|
|
100
104
|
| Recipient | Files the lead lists under `## Inputs` |
|
|
101
105
|
|---|---|
|
|
102
|
-
|
|
|
106
|
+
| Any rostered analysis worker | `analysis-packet.md` as primary input; for `final-verification`, no source/fallback list is copied into the prompt |
|
|
103
107
|
| Report writer worker (Phase 6) | task-brief, analysis-profile, analysis-material, reference-expectations, clarification-response (if carry-in), **plus** the instruction-set-local `final-report-template.md` (phase-stripped) and `final-report-schema.json` (per-task-type excerpt) — NOT the full `templates/reports/...` / `schemas/...` sources |
|
|
104
108
|
| Reverify dispatches | none — the lead provides only the items to reverify |
|
|
105
109
|
|
|
@@ -37,7 +37,7 @@ are collected and convergence finished. Phase 1-5 do not need it.
|
|
|
37
37
|
evidence and add one `recommendedNextSteps` item with environment
|
|
38
38
|
prerequisites and the expected `QA-RESULT`. This is user-owned verification;
|
|
39
39
|
do not route back or fail implementation solely for this advisory.
|
|
40
|
-
- **Routing recommendation
|
|
40
|
+
- **Routing recommendation**: `implementation.routingRecommendation` is an **object** with exactly two fields — `target`, one of `final-verification`, `error-analysis`, `implementation-planning`, `implementation`, and `rationale`, one or two sentences on why that target and nothing else. It is not a prose note: Phase 7 projects `workflow.nextRecommendedPhase` from `target` alone, so a phase named only in the prose does not route the task. Pick `final-verification` when this stage's plan items landed and validation passed; `error-analysis` when a failure's cause is not understood; `implementation-planning` when the approved plan itself no longer fits the evidence; `implementation` when work remains inside this stage and the next run is a fix run. **Enforced:** `schemas/final-report-v2.0.schema.json` rejects a `target` outside the enum, a missing `rationale`, and a string in place of the object.
|
|
41
41
|
- **Follow-up tasks (Section 4 of the final report)**: every item discovered during this run that was *not* delivered MUST appear in the final report's `## 4. Follow-up Tasks` table with a concrete `Origin`, `New Task ID`, `Suggested task-type`, `Scope`, and `Reason / Why deferred`. Sources include: out-of-scope discoveries that the executor consciously chose not to fold into this run, verifier concerns the executor declined to fix in-place, scope-boundary items from the approved plan that turned out to need their own ticket, and any unresolved `## 1. Clarification Items` row carried over from the approved plan (`Status` ∈ `{open, answered}` at approval time). An empty section is acceptable but only when expressed as the single line `- No follow-up tasks.` — silence is treated as a contract violation. Rows with `Auto-spawn? = yes` will be materialised by `scripts/okstra-spawn-followups.py` in Phase 7; rows with `Auto-spawn? = no` MUST also appear in `Section 3. Recommended Next Steps` so the user knows to act manually.
|
|
42
42
|
|
|
43
43
|
## Self-review pass before finalising the report (the Okstra lead runs this; do not delegate it)
|
|
@@ -7,11 +7,8 @@ until Phase 5 ends, then drop from active context for Phase 6/7.
|
|
|
7
7
|
|
|
8
8
|
# Implementation profile — Executor sidecar
|
|
9
9
|
|
|
10
|
-
> **When to read**: lead reads this file ONCE at the start of Phase 5 (after Stage Map parse, before issuing the Executor's first `Edit` / `Write`). The body governs ONLY the Executor role's behaviour. Verifier / report-writer behaviour lives in sibling sidecars.
|
|
11
|
-
|
|
12
10
|
## Executor role binding (carried over from the thin core)
|
|
13
11
|
|
|
14
|
-
- **Executor dispatch labelling.** The core functional role label is `<provider>-executor` (e.g. `codex-executor`). Provider, role, and model identity are owned by `prompts/lead/okstra-lead-contract.md` "Model assignments"; the selected runtime adapter owns provider-native dispatch-label mapping (including any `name` / `**Pane role:**` fields) and token-attribution wiring under its "Semantic operation mapping". This functional label is NOT what the run's PROGRESS checkpoints carry: `phase-4-dispatch` / `phase-5-collect` name the roster role team-state records (`Codex worker`), because that is the entry the Phase 7 conformance check matches them against.
|
|
15
12
|
- The `Executor` (bound in `implementation.md` thin core) is the **only worker permitted to mutate project files**. All other workers run read-only. A `runner=native-session` executor uses the selected host adapter's native edit and command primitives. A `runner=cli-wrapper` executor mutates files inside its provider CLI's auto-edit mode. The safety rules in this sidecar apply identically to both runners.
|
|
16
13
|
- When the thin core's Task worktree block resolves status to `created` or `reused`, the Executor MUST run every Edit / Write / build / test / commit command with the worktree path as cwd. Treat it as `project_root` for the duration of this run. Do NOT mutate the caller's original checkout. Do NOT `cd` out of the worktree to reach files. If a file outside the worktree is genuinely needed, treat it as a planning gap: record it in `Out-of-plan edits` and continue.
|
|
17
14
|
- **How to set the working directory**: every command and native edit MUST target `{{EXECUTOR_WORKTREE_PATH}}`, never the lead session's original project directory. The selected runtime adapter owns the exact native command syntax. Provider CLI wrappers inject the worktree at the CLI layer. For tools that accept an explicit working-directory flag (`git -C <path>`, `cargo --manifest-path`, `pytest --rootdir`), prefer that form.
|
|
@@ -5,14 +5,10 @@ at Phase 5, BEFORE constructing the verifier worker dispatch prompts.
|
|
|
5
5
|
|
|
6
6
|
# Implementation profile — Verifier sidecar
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
## Verifier independence
|
|
9
9
|
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
- **Verifier dispatch labelling.** The core functional role label is `<provider>-verifier` (here, and identically in `final-verification`). Provider, role, and model identity are owned by `prompts/lead/okstra-lead-contract.md` "Model assignments"; the selected runtime adapter owns provider-native dispatch-label mapping (including any `name` / `**Pane role:**` fields) and token-attribution wiring under its "Semantic operation mapping". This functional label is NOT what the run's PROGRESS checkpoints carry: `phase-4-dispatch` / `phase-5-collect` name the roster role team-state records (`Claude worker`, `Codex worker`), because that is the entry the Phase 7 conformance check matches them against.
|
|
13
|
-
- The verifier slots are `Claude verifier` and `Codex verifier`, plus `Antigravity verifier` **only when `antigravity` is in the resolved `--workers` roster**. Every verifier in the resolved roster is dispatched except the one whose worker ID holds the executor role this run: that ID materializes as the executor on every dispatch (`scripts/okstra_ctl/worker_prompt_policy.py`), so the executor's own provider has no separate verifier session in the current plumbing — a follow-up design item. Independence still holds where it counts: every verdict comes from a fresh CLI session with no shared context, never from the session that wrote the diff. Verifiers MUST NOT call Edit, Write, or any Bash command that mutates files outside the run's artifact directories. If a verifier wants a fix, it records the recommendation in its worker result; it does not apply the fix itself.
|
|
14
|
-
- Session isolation — not model-variant divergence — is the primary self-review safeguard: each verifier is a separate CLI invocation with its own context window, so a verifier reusing the executor's model variant is acceptable. Different model variants (e.g. executor=opus / Claude verifier=sonnet) remain recommended when available.
|
|
15
|
-
- Phase-specific model defaults override the shared defaults: `Claude verifier`=`opus`, `Codex verifier`=`gpt-5.6-sol`, `Antigravity verifier`=`gemini-3.1-pro` (only when present in the roster). The `Executor`'s model is taken from the provider-specific worker model corresponding to `--executor`: claude→`--claude-model` (default `opus`), codex→`--codex-model` (default `gpt-5.6-sol`), antigravity→`--antigravity-model` (default `gemini-3.1-pro`).
|
|
10
|
+
- Every verdict comes from a fresh session with no shared context, never from the session that wrote the diff. Verifiers MUST NOT call Edit, Write, or any Bash command that mutates files outside the run's artifact directories. If a verifier wants a fix, it records the recommendation in its worker result; it does not apply the fix itself.
|
|
11
|
+
- Session isolation is the primary self-review safeguard: each verifier is a separate invocation with its own context window. Reusing the executor's model is acceptable. The model comes from the run's stored assignment.
|
|
16
12
|
- Verifiers read from the SAME working tree path the Executor used so they observe the exact diff the Executor produced. Verifiers remain strictly read-only there.
|
|
17
13
|
|
|
18
14
|
## Verifier QA duties (independent re-run mandate)
|
|
@@ -3,15 +3,19 @@
|
|
|
3
3
|
```yaml
|
|
4
4
|
roles:
|
|
5
5
|
- role: analyser
|
|
6
|
-
|
|
7
|
-
|
|
6
|
+
min: 2
|
|
7
|
+
recommended: 2
|
|
8
|
+
max: 5
|
|
8
9
|
duty: analysis-worker
|
|
9
|
-
minDistinctProviders: 2
|
|
10
10
|
- role: report-writer
|
|
11
|
-
|
|
11
|
+
min: 1
|
|
12
|
+
recommended: 1
|
|
13
|
+
max: 1
|
|
12
14
|
duty: report-writer
|
|
13
15
|
- role: verifier
|
|
14
|
-
|
|
16
|
+
min: 0
|
|
17
|
+
recommended: 0
|
|
18
|
+
max: 0
|
|
15
19
|
duty: reverification-worker
|
|
16
20
|
dynamic: true
|
|
17
21
|
```
|
|
@@ -3,19 +3,24 @@
|
|
|
3
3
|
```yaml
|
|
4
4
|
roles:
|
|
5
5
|
- role: analyser
|
|
6
|
-
|
|
7
|
-
|
|
6
|
+
min: 2
|
|
7
|
+
recommended: 2
|
|
8
|
+
max: 5
|
|
8
9
|
duty: diagnosis-worker
|
|
9
|
-
minDistinctProviders: 2
|
|
10
10
|
- role: critic
|
|
11
|
-
|
|
12
|
-
|
|
11
|
+
min: 0
|
|
12
|
+
recommended: 0
|
|
13
|
+
max: 1
|
|
13
14
|
duty: scope-critic
|
|
14
15
|
- role: report-writer
|
|
15
|
-
|
|
16
|
+
min: 1
|
|
17
|
+
recommended: 1
|
|
18
|
+
max: 1
|
|
16
19
|
duty: report-writer
|
|
17
20
|
- role: verifier
|
|
18
|
-
|
|
21
|
+
min: 0
|
|
22
|
+
recommended: 0
|
|
23
|
+
max: 0
|
|
19
24
|
duty: reverification-worker
|
|
20
25
|
dynamic: true
|
|
21
26
|
```
|
|
@@ -37,13 +42,14 @@ roles:
|
|
|
37
42
|
- any `intent-inference` augmentation that re-characterises the symptom (e.g. classifying a vague reporter phrase like "it sometimes doesn't work" as "intermittent failure on a specific code path") is a **hypothesis**, not a confirmed symptom. If `[CONFIRMED …]` appears on the matching `intent-check:` row, treat that confirmation as the symptom. Otherwise follow the precondition's `skipped` branch above and keep the inference labelled as a hypothesis in the root-cause analysis.
|
|
38
43
|
- `conversion-block:` rows mean the brief could not map a reporter statement to project vocabulary; never invent the missing mapping in this phase.
|
|
39
44
|
- Worker diagnosis procedure:
|
|
45
|
+
- **Ticket Tagging.** Tag every section 1–5 item with its related ticket. Use `Issue / Ticket`, fall back to Task ID, then `unknown`; comma-separate multiple tickets.
|
|
40
46
|
- **Symptom lock:** state the reporter's symptom verbatim, then translate it into one observable failure condition. If no observable condition can be derived from the brief, record that gap as the first blocker instead of guessing.
|
|
41
47
|
- **Pass condition lock:** the brief's `EB-NNN` items state what the system must do once the defect is gone. State, per id, the observation that would show the symptom resolved — this is the upper bound on the fix. A diagnosis that leaves "how far do we fix this" open is what lets the later plan expand. Record each as an `endStateCoverage` row whose `coveredBy` names the root-cause candidate or next diagnostic that accounts for it. **Enforced:** `validators/validate-run.py` `_validate_end_state_coverage`.
|
|
42
48
|
- **Reproduction status:** classify the run as `reproduced`, `not-reproduced`, or `blocked-before-repro`. Cite the command/log/file evidence used. If no command can be run safely in this phase, explain the read-only evidence path and the exact material needed next.
|
|
43
49
|
- **Falsifiable cause candidates:** every root-cause candidate must include supporting evidence, the strongest falsifying evidence checked, confidence, and the next diagnostic action that would disprove it. A candidate that cannot be falsified is too vague for this phase.
|
|
44
50
|
- **Graph-aware scope:** a graph edge can explain ordering or duplication, but it is not proof of cause by itself. Cite code/log evidence before claiming an upstream related task caused the current symptom.
|
|
45
51
|
- **Sharp next diagnostic:** end with the single highest-value diagnostic command, log capture, or file inspection that should happen next, plus the expected signal that would confirm or reject the leading cause.
|
|
46
|
-
- **Fix-design boundary:** do not design the implementation fix beyond what is necessary to validate the cause. If the cause is credible,
|
|
52
|
+
- **Fix-design boundary:** do not design the implementation fix beyond what is necessary to validate the cause. If the cause is credible, recommend `implementation-option-selection` with the verified evidence; if the cause is still unclear, recommend another `error-analysis` run with the next diagnostic. The lead's Phase Routing settles the next phase.
|
|
47
53
|
- Structured diagnosis and routing contract:
|
|
48
54
|
- `errorAnalysis` is the source of truth for reproduction status, `EA-NNN` cause candidates, the sharp next diagnostic, and the next route.
|
|
49
55
|
- A route to `implementation-option-selection` requires a credible leading cause referenced by `routing.leadingCauseId` and `begin-option-selection` as the direction. A route back to `error-analysis` requires the sharp next diagnostic and `continue-investigation` as the direction.
|
|
@@ -3,15 +3,19 @@
|
|
|
3
3
|
```yaml
|
|
4
4
|
roles:
|
|
5
5
|
- role: analyser
|
|
6
|
-
|
|
7
|
-
|
|
6
|
+
min: 2
|
|
7
|
+
recommended: 2
|
|
8
|
+
max: 5
|
|
8
9
|
duty: analysis-worker
|
|
9
|
-
minDistinctProviders: 2
|
|
10
10
|
- role: report-writer
|
|
11
|
-
|
|
11
|
+
min: 1
|
|
12
|
+
recommended: 1
|
|
13
|
+
max: 1
|
|
12
14
|
duty: report-writer
|
|
13
15
|
- role: verifier
|
|
14
|
-
|
|
16
|
+
min: 0
|
|
17
|
+
recommended: 0
|
|
18
|
+
max: 0
|
|
15
19
|
duty: reverification-worker
|
|
16
20
|
dynamic: true
|
|
17
21
|
```
|