okstra 0.164.0 → 0.165.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +1 -1
- package/docs/architecture.md +12 -8
- package/docs/cli.md +7 -3
- package/docs/for-ai/README.md +2 -2
- package/docs/for-ai/skills/okstra-inspect.md +2 -2
- package/docs/for-ai/skills/okstra-user-response.md +2 -2
- package/docs/project-structure-overview.md +15 -9
- package/package.json +1 -1
- package/runtime/BUILD.json +2 -2
- package/runtime/agents/workers/antigravity-worker.md +9 -7
- package/runtime/agents/workers/codex-worker.md +9 -7
- package/runtime/agents/workers/grok-worker.md +6 -4
- package/runtime/agents/workers/kimi-worker.md +6 -4
- package/runtime/bin/okstra-antigravity-exec.sh +1 -340
- package/runtime/bin/okstra-claude-exec.sh +1 -178
- package/runtime/bin/okstra-codex-exec.sh +1 -467
- package/runtime/bin/okstra-provider-exec.py +165 -190
- package/runtime/bin/okstra-trace-cleanup.sh +14 -7
- package/runtime/bin/okstra-wrapper-status.py +26 -19
- package/runtime/prompts/lead/adapters/cmux.md +1 -1
- package/runtime/prompts/lead/convergence.md +36 -8
- package/runtime/prompts/lead/okstra-lead-contract.md +23 -1
- package/runtime/prompts/lead/plan-body-verification.md +9 -1
- package/runtime/prompts/lead/report-writer.md +1 -0
- package/runtime/prompts/lead/team-contract.md +3 -3
- package/runtime/prompts/profiles/_common-contract.md +9 -1
- package/runtime/prompts/profiles/_coverage-critic.md +1 -1
- package/runtime/prompts/profiles/_implementation-diff-review.md +3 -1
- package/runtime/prompts/profiles/_implementation-self-check.md +1 -1
- package/runtime/prompts/profiles/_implementation-verifier.md +3 -1
- package/runtime/prompts/profiles/implementation-planning.md +5 -3
- package/runtime/python/okstra_ctl/adapters/hosts/claude-code/relay.md +1 -1
- package/runtime/python/okstra_ctl/adapters/hosts/external/relay.md +1 -1
- package/runtime/python/okstra_ctl/adapters/providers/antigravity/adapter.py +148 -0
- package/runtime/python/okstra_ctl/adapters/providers/claude/adapter.py +55 -0
- package/runtime/python/okstra_ctl/adapters/providers/codex/adapter.py +41 -0
- package/runtime/python/okstra_ctl/adapters/providers/grok/adapter.py +44 -0
- package/runtime/python/okstra_ctl/adapters/providers/kimi/adapter.py +42 -0
- package/runtime/python/okstra_ctl/dispatch_core.py +5 -1
- package/runtime/python/okstra_ctl/dispatch_state.py +10 -0
- package/runtime/python/okstra_ctl/domain/provider.py +5 -1
- package/runtime/python/okstra_ctl/domain/worker_exec.py +102 -0
- package/runtime/python/okstra_ctl/domain/worker_role.py +34 -0
- package/runtime/python/okstra_ctl/domain/worker_stream.py +261 -0
- package/runtime/python/okstra_ctl/incremental_scope.py +16 -4
- package/runtime/python/okstra_ctl/report_html/common.py +71 -25
- package/runtime/python/okstra_ctl/report_html/models.py +5 -0
- package/runtime/python/okstra_ctl/report_html/render.py +1 -1
- package/runtime/python/okstra_ctl/report_html/run_usage.py +19 -0
- package/runtime/python/okstra_ctl/report_html/view_models/implementation_planning.py +14 -0
- package/runtime/python/okstra_ctl/report_views.py +44 -16
- package/runtime/python/okstra_ctl/stage_citations.py +52 -15
- package/runtime/python/okstra_ctl/user_response.py +45 -29
- package/runtime/python/okstra_ctl/wizard.py +13 -9
- package/runtime/python/okstra_ctl/worker_prompt_policy.py +10 -3
- package/runtime/python/okstra_ctl/worker_request.py +140 -0
- package/runtime/python/okstra_ctl/worker_runner.py +622 -0
- package/runtime/python/okstra_token_usage/collect.py +8 -1
- package/runtime/python/okstra_token_usage/report.py +42 -0
- package/runtime/python/okstra_token_usage/task_totals.py +88 -0
- package/runtime/schemas/final-report-v1.0.schema.json +70 -0
- package/runtime/schemas/final-report-v2.0.schema.json +90 -0
- package/runtime/skills/okstra-inspect/SKILL.md +1 -2
- package/runtime/skills/okstra-inspect/facets/logs.md +5 -5
- package/runtime/skills/okstra-inspect/facets/run-audit.md +3 -3
- package/runtime/skills/okstra-run/SKILL.md +1 -1
- package/runtime/skills/okstra-user-response/SKILL.md +15 -5
- package/runtime/templates/report-writer-prompt-preamble.md +1 -0
- package/runtime/templates/reports/html/assets/base.css +8 -4
- package/runtime/templates/reports/html/base.template.html +12 -6
- package/runtime/templates/reports/html/i18n/en.json +29 -6
- package/runtime/templates/reports/html/i18n/ko.json +29 -6
- package/runtime/templates/reports/html/macros/forms.html +9 -3
- package/runtime/templates/reports/html/tasks/implementation-planning.template.html +14 -19
- package/runtime/templates/reports/report.js +59 -26
- package/runtime/templates/reports/user-response.template.md +12 -8
- package/runtime/validators/validate-run.py +88 -7
- package/runtime/validators/validate_session_conformance.py +62 -1
- package/src/cli-registry.mjs +0 -7
- package/runtime/bin/okstra-wrapper-agy-stream.py +0 -61
- package/runtime/python/okstra_ctl/error_issue.py +0 -640
- package/runtime/python/okstra_ctl/issue_signals.py +0 -186
- package/runtime/skills/okstra-inspect/facets/error-issue.md +0 -77
- package/src/commands/inspect/error-issue.mjs +0 -27
|
@@ -1,121 +1,184 @@
|
|
|
1
1
|
#!/usr/bin/env python3
|
|
2
|
-
"""Run an external LLM CLI with the shared okstra wrapper contract.
|
|
2
|
+
"""Run an external LLM CLI with the shared okstra wrapper contract.
|
|
3
|
+
|
|
4
|
+
Argument parsing and provider lookup only — the run itself belongs to
|
|
5
|
+
``okstra_ctl.worker_runner``, which every provider shares. What stays here is
|
|
6
|
+
the request the provider strategies then take on trust: resolved paths, the
|
|
7
|
+
write scope in the order the CLIs are told it, and the role's idle budget.
|
|
8
|
+
"""
|
|
3
9
|
from __future__ import annotations
|
|
4
10
|
|
|
5
|
-
import json
|
|
6
11
|
import os
|
|
7
|
-
import selectors
|
|
8
12
|
import shutil
|
|
9
|
-
import signal
|
|
10
|
-
import subprocess
|
|
11
13
|
import sys
|
|
12
|
-
import time
|
|
13
14
|
from dataclasses import dataclass
|
|
14
15
|
from pathlib import Path
|
|
15
|
-
from typing import
|
|
16
|
+
from typing import Any, Mapping
|
|
17
|
+
|
|
18
|
+
_HERE = Path(__file__).resolve().parent
|
|
19
|
+
# ``okstra_ctl`` sits beside this file in the repo (``scripts/``) but under
|
|
20
|
+
# ``~/.okstra/lib/python/`` once installed, and the four-line shell entrypoints
|
|
21
|
+
# that exec this script set no PYTHONPATH. Offer both, repo first.
|
|
22
|
+
_HOME_LIB = (
|
|
23
|
+
Path(os.environ.get("OKSTRA_HOME", str(Path.home() / ".okstra"))) / "lib" / "python"
|
|
24
|
+
)
|
|
25
|
+
sys.path.insert(0, str(_HERE))
|
|
26
|
+
if _HOME_LIB.is_dir() and str(_HOME_LIB) not in sys.path:
|
|
27
|
+
sys.path.append(str(_HOME_LIB))
|
|
28
|
+
|
|
29
|
+
from okstra_ctl.domain.provider import ProviderSpec, UnknownProviderError # noqa: E402
|
|
30
|
+
from okstra_ctl.domain.worker_exec import ( # noqa: E402
|
|
31
|
+
ExecutionStrategy,
|
|
32
|
+
WorkerExecRequest,
|
|
33
|
+
)
|
|
34
|
+
from okstra_ctl.registry.provider_registry import ( # noqa: E402
|
|
35
|
+
default_provider_registry,
|
|
36
|
+
)
|
|
37
|
+
from okstra_ctl.worker_request import build_request, idle_timeout # noqa: E402
|
|
38
|
+
from okstra_ctl.worker_runner import LIVE, QUIET, run_worker # noqa: E402
|
|
39
|
+
|
|
40
|
+
_USAGE = (
|
|
41
|
+
"usage: okstra-provider-exec.py <provider> <project-root> "
|
|
42
|
+
"<model-execution-value> <prompt-path> [worktree-path] [role] "
|
|
43
|
+
"[idle-timeout-seconds] [--presentation live|quiet]"
|
|
44
|
+
)
|
|
45
|
+
|
|
46
|
+
_PRESENTATION_FLAG = "--presentation"
|
|
47
|
+
_PRESENTATIONS = (LIVE, QUIET)
|
|
48
|
+
|
|
49
|
+
|
|
50
|
+
class PreflightError(Exception):
|
|
51
|
+
def __init__(self, exit_code: int, message: str) -> None:
|
|
52
|
+
super().__init__(message)
|
|
53
|
+
self.exit_code = exit_code
|
|
16
54
|
|
|
17
55
|
|
|
18
56
|
@dataclass(frozen=True)
|
|
19
|
-
class
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
57
|
+
class Invocation:
|
|
58
|
+
strategy: ExecutionStrategy
|
|
59
|
+
request: WorkerExecRequest
|
|
60
|
+
presentation: str
|
|
61
|
+
log_path: Path
|
|
62
|
+
status_path: Path
|
|
63
|
+
status_extra: Mapping[str, Any]
|
|
64
|
+
|
|
65
|
+
|
|
66
|
+
def parse_invocation(argv: list[str]) -> Invocation:
|
|
67
|
+
"""Resolve the wrapper's positional contract into one runnable dispatch."""
|
|
68
|
+
positional, presentation = _take_presentation(argv)
|
|
69
|
+
if not 4 <= len(positional) <= 7:
|
|
70
|
+
raise PreflightError(64, _USAGE)
|
|
71
|
+
provider_id, project_root_raw, model, prompt_raw = positional[:4]
|
|
72
|
+
worktree_raw = positional[4] if len(positional) >= 5 else ""
|
|
73
|
+
role = positional[5] if len(positional) >= 6 and positional[5] else "worker"
|
|
74
|
+
timeout_raw = positional[6] if len(positional) >= 7 and positional[6] else ""
|
|
75
|
+
|
|
76
|
+
spec = _provider_spec(provider_id)
|
|
77
|
+
project_root = _existing_dir(project_root_raw, 65, "project-root")
|
|
78
|
+
if not model:
|
|
79
|
+
raise PreflightError(66, "model-execution-value is empty")
|
|
80
|
+
prompt_path = _existing_file(prompt_raw, 67, "prompt-path")
|
|
81
|
+
idle_timeout_seconds = _idle_timeout(timeout_raw, role)
|
|
82
|
+
worktree = (
|
|
83
|
+
_existing_dir(worktree_raw, 68, "worktree-path") if worktree_raw else None
|
|
84
|
+
)
|
|
23
85
|
|
|
86
|
+
request = build_request(
|
|
87
|
+
prompt_text=prompt_path.read_text(encoding="utf-8"),
|
|
88
|
+
model=model,
|
|
89
|
+
project_root=project_root,
|
|
90
|
+
worktree_path=worktree,
|
|
91
|
+
role=role,
|
|
92
|
+
idle_timeout_seconds=idle_timeout_seconds,
|
|
93
|
+
)
|
|
94
|
+
strategy = spec.exec_strategy
|
|
95
|
+
_check_command(strategy, request)
|
|
96
|
+
return Invocation(
|
|
97
|
+
strategy=strategy,
|
|
98
|
+
request=request,
|
|
99
|
+
presentation=presentation,
|
|
100
|
+
log_path=_log_path(prompt_path),
|
|
101
|
+
status_path=Path(f"{prompt_path}.status.json"),
|
|
102
|
+
status_extra={"wrapper": spec.wrapper, "role": role},
|
|
103
|
+
)
|
|
24
104
|
|
|
25
|
-
def _grok_args(prompt: str, model: str, cwd: str) -> list[str]:
|
|
26
|
-
return [
|
|
27
|
-
"grok",
|
|
28
|
-
"-p",
|
|
29
|
-
prompt,
|
|
30
|
-
"-m",
|
|
31
|
-
model,
|
|
32
|
-
"--output-format",
|
|
33
|
-
"streaming-json",
|
|
34
|
-
"--cwd",
|
|
35
|
-
cwd,
|
|
36
|
-
]
|
|
37
105
|
|
|
106
|
+
def _take_presentation(argv: list[str]) -> tuple[list[str], str]:
|
|
107
|
+
"""Split the one flag out of an otherwise positional argv.
|
|
108
|
+
|
|
109
|
+
Defaults to ``quiet``. ``live`` is only ever right where a screen was
|
|
110
|
+
declared, and the only callers that can declare one are the pane backends —
|
|
111
|
+
which pass the flag explicitly. Defaulting the other way assumed a screen
|
|
112
|
+
that a subagent dispatch does not have, and sent every worker's progress
|
|
113
|
+
into its caller's context window instead.
|
|
114
|
+
"""
|
|
115
|
+
positional: list[str] = []
|
|
116
|
+
presentation = QUIET
|
|
117
|
+
index = 0
|
|
118
|
+
while index < len(argv):
|
|
119
|
+
if argv[index] != _PRESENTATION_FLAG:
|
|
120
|
+
positional.append(argv[index])
|
|
121
|
+
index += 1
|
|
122
|
+
continue
|
|
123
|
+
if index + 1 >= len(argv):
|
|
124
|
+
raise PreflightError(64, f"{_PRESENTATION_FLAG} needs a value: {_USAGE}")
|
|
125
|
+
presentation = argv[index + 1]
|
|
126
|
+
index += 2
|
|
127
|
+
if presentation not in _PRESENTATIONS:
|
|
128
|
+
allowed = " | ".join(_PRESENTATIONS)
|
|
129
|
+
raise PreflightError(
|
|
130
|
+
64, f"unsupported presentation {presentation!r}. Allowed values: {allowed}"
|
|
131
|
+
)
|
|
132
|
+
return positional, presentation
|
|
38
133
|
|
|
39
|
-
def _kimi_args(prompt: str, model: str, _cwd: str) -> list[str]:
|
|
40
|
-
return ["kimi", "-p", prompt, "-m", model, "--output-format", "stream-json"]
|
|
41
134
|
|
|
135
|
+
def _provider_spec(provider_id: str) -> ProviderSpec:
|
|
136
|
+
try:
|
|
137
|
+
spec = default_provider_registry().resolve(provider_id)
|
|
138
|
+
except UnknownProviderError as exc:
|
|
139
|
+
raise PreflightError(64, str(exc)) from exc
|
|
140
|
+
if spec.exec_strategy is None:
|
|
141
|
+
raise PreflightError(
|
|
142
|
+
64, f"provider {provider_id!r} has no execution strategy to run"
|
|
143
|
+
)
|
|
144
|
+
return spec
|
|
42
145
|
|
|
43
|
-
PROVIDERS = {
|
|
44
|
-
"grok": ProviderCommand("grok", "okstra-grok-exec.sh", _grok_args),
|
|
45
|
-
"kimi": ProviderCommand("kimi", "okstra-kimi-exec.sh", _kimi_args),
|
|
46
|
-
}
|
|
47
146
|
|
|
147
|
+
def _existing_dir(raw: str, exit_code: int, label: str) -> Path:
|
|
148
|
+
"""Existence only — `build_request` owns the resolving."""
|
|
149
|
+
path = Path(raw) if raw else None
|
|
150
|
+
if path is None or not path.is_dir():
|
|
151
|
+
raise PreflightError(
|
|
152
|
+
exit_code, f"{label} is missing or not a directory: {raw!r}"
|
|
153
|
+
)
|
|
154
|
+
return path
|
|
48
155
|
|
|
49
|
-
@dataclass(frozen=True)
|
|
50
|
-
class Invocation:
|
|
51
|
-
provider: ProviderCommand
|
|
52
|
-
project_root: Path
|
|
53
|
-
model: str
|
|
54
|
-
prompt_path: Path
|
|
55
|
-
execution_root: Path
|
|
56
|
-
role: str
|
|
57
|
-
idle_timeout_seconds: int
|
|
58
156
|
|
|
157
|
+
def _existing_file(raw: str, exit_code: int, label: str) -> Path:
|
|
158
|
+
path = Path(raw) if raw else None
|
|
159
|
+
if path is None or not path.is_file():
|
|
160
|
+
raise PreflightError(exit_code, f"{label} is missing or not a file: {raw!r}")
|
|
161
|
+
return path.resolve()
|
|
59
162
|
|
|
60
|
-
class PreflightError(Exception):
|
|
61
|
-
def __init__(self, exit_code: int, message: str) -> None:
|
|
62
|
-
super().__init__(message)
|
|
63
|
-
self.exit_code = exit_code
|
|
64
163
|
|
|
164
|
+
def _idle_timeout(raw: str, role: str) -> int:
|
|
165
|
+
try:
|
|
166
|
+
return idle_timeout(raw, role)
|
|
167
|
+
except ValueError as exc:
|
|
168
|
+
raise PreflightError(69, str(exc)) from exc
|
|
65
169
|
|
|
66
|
-
def _parse_invocation(argv: list[str]) -> Invocation:
|
|
67
|
-
if len(argv) < 4 or len(argv) > 7:
|
|
68
|
-
raise PreflightError(
|
|
69
|
-
64,
|
|
70
|
-
"usage: okstra-provider-exec.py <provider> <project-root> <model-execution-value> "
|
|
71
|
-
"<prompt-path> [worktree-path] [role] [idle-timeout-seconds]",
|
|
72
|
-
)
|
|
73
|
-
provider_id, project_root_raw, model, prompt_raw = argv[:4]
|
|
74
|
-
provider = PROVIDERS.get(provider_id)
|
|
75
|
-
if provider is None:
|
|
76
|
-
raise PreflightError(64, f"unsupported provider: {provider_id}")
|
|
77
|
-
worktree_raw = argv[4] if len(argv) >= 5 else ""
|
|
78
|
-
role = argv[5] if len(argv) >= 6 and argv[5] else "worker"
|
|
79
|
-
default_timeout = 1500 if role in {"executor", "verifier"} else 600
|
|
80
|
-
timeout_raw = argv[6] if len(argv) >= 7 else str(default_timeout)
|
|
81
|
-
return _validate_invocation(
|
|
82
|
-
provider, project_root_raw, model, prompt_raw, worktree_raw, role, timeout_raw
|
|
83
|
-
)
|
|
84
170
|
|
|
171
|
+
def _check_command(strategy: ExecutionStrategy, request: WorkerExecRequest) -> None:
|
|
172
|
+
"""Refuse a missing CLI before the run leaves any artifact behind.
|
|
85
173
|
|
|
86
|
-
|
|
87
|
-
provider
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
) -> Invocation:
|
|
95
|
-
project_root = Path(project_root_raw)
|
|
96
|
-
prompt_path = Path(prompt_raw)
|
|
97
|
-
if not project_root_raw or not project_root.is_dir():
|
|
98
|
-
raise PreflightError(65, f"project-root is missing or not a directory: {project_root_raw!r}")
|
|
99
|
-
if not model:
|
|
100
|
-
raise PreflightError(66, "model-execution-value is empty")
|
|
101
|
-
if not prompt_raw or not prompt_path.is_file():
|
|
102
|
-
raise PreflightError(67, f"prompt-path is missing or not a file: {prompt_raw!r}")
|
|
103
|
-
if not timeout_raw.isdigit():
|
|
104
|
-
raise PreflightError(69, f"idle-timeout-seconds must be a non-negative integer: {timeout_raw!r}")
|
|
105
|
-
execution_root = Path(worktree_raw) if worktree_raw else project_root
|
|
106
|
-
if worktree_raw and not execution_root.is_dir():
|
|
107
|
-
raise PreflightError(68, f"worktree-path was provided but is not a directory: {worktree_raw!r}")
|
|
108
|
-
if shutil.which(provider.binary) is None:
|
|
109
|
-
raise PreflightError(127, f"{provider.binary} CLI is not installed on PATH")
|
|
110
|
-
return Invocation(
|
|
111
|
-
provider=provider,
|
|
112
|
-
project_root=project_root.resolve(),
|
|
113
|
-
model=model,
|
|
114
|
-
prompt_path=prompt_path.resolve(),
|
|
115
|
-
execution_root=execution_root.resolve(),
|
|
116
|
-
role=role,
|
|
117
|
-
idle_timeout_seconds=int(timeout_raw),
|
|
118
|
-
)
|
|
174
|
+
Building the command is the only truthful way to learn which binary this
|
|
175
|
+
provider runs, and it is a pure call. The check stays out of the runner so a
|
|
176
|
+
refused dispatch leaves no `started` status sidecar for the liveness probe to
|
|
177
|
+
read as a worker that launched.
|
|
178
|
+
"""
|
|
179
|
+
binary = strategy.build_command(request).argv[0]
|
|
180
|
+
if shutil.which(binary) is None:
|
|
181
|
+
raise PreflightError(127, f"{binary} CLI is not installed on PATH")
|
|
119
182
|
|
|
120
183
|
|
|
121
184
|
def _log_path(prompt_path: Path) -> Path:
|
|
@@ -124,105 +187,17 @@ def _log_path(prompt_path: Path) -> Path:
|
|
|
124
187
|
return Path(f"{prompt_path}.log")
|
|
125
188
|
|
|
126
189
|
|
|
127
|
-
def _write_status(path: Path, status: dict[str, object]) -> None:
|
|
128
|
-
temporary = Path(f"{path}.tmp")
|
|
129
|
-
temporary.write_text(json.dumps(status, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
|
130
|
-
os.replace(temporary, path)
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
def _terminate_process(process: subprocess.Popen[bytes]) -> None:
|
|
134
|
-
try:
|
|
135
|
-
os.killpg(process.pid, signal.SIGTERM)
|
|
136
|
-
except ProcessLookupError:
|
|
137
|
-
return
|
|
138
|
-
try:
|
|
139
|
-
process.wait(timeout=5)
|
|
140
|
-
except subprocess.TimeoutExpired:
|
|
141
|
-
try:
|
|
142
|
-
os.killpg(process.pid, signal.SIGKILL)
|
|
143
|
-
except ProcessLookupError:
|
|
144
|
-
pass
|
|
145
|
-
process.wait()
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
def _stream_process(
|
|
149
|
-
process: subprocess.Popen[bytes], log_file, idle_timeout_seconds: int
|
|
150
|
-
) -> tuple[int, bool, int]:
|
|
151
|
-
selector = selectors.DefaultSelector()
|
|
152
|
-
assert process.stdout is not None
|
|
153
|
-
selector.register(process.stdout, selectors.EVENT_READ)
|
|
154
|
-
last_output = time.monotonic()
|
|
155
|
-
timed_out = False
|
|
156
|
-
idle_seconds = 0
|
|
157
|
-
while selector.get_map():
|
|
158
|
-
for key, _ in selector.select(timeout=0.25):
|
|
159
|
-
chunk = os.read(key.fd, 8192)
|
|
160
|
-
if not chunk:
|
|
161
|
-
selector.unregister(key.fileobj)
|
|
162
|
-
continue
|
|
163
|
-
last_output = time.monotonic()
|
|
164
|
-
sys.stdout.buffer.write(chunk)
|
|
165
|
-
sys.stdout.buffer.flush()
|
|
166
|
-
log_file.write(chunk)
|
|
167
|
-
log_file.flush()
|
|
168
|
-
idle_seconds = int(time.monotonic() - last_output)
|
|
169
|
-
if idle_timeout_seconds and idle_seconds >= idle_timeout_seconds and process.poll() is None:
|
|
170
|
-
timed_out = True
|
|
171
|
-
_terminate_process(process)
|
|
172
|
-
exit_code = process.wait()
|
|
173
|
-
return (124 if timed_out else exit_code), timed_out, idle_seconds
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
def _run(invocation: Invocation) -> int:
|
|
177
|
-
prompt = invocation.prompt_path.read_text(encoding="utf-8")
|
|
178
|
-
command = invocation.provider.build_args(prompt, invocation.model, str(invocation.execution_root))
|
|
179
|
-
status_path = Path(f"{invocation.prompt_path}.status.json")
|
|
180
|
-
log_path = _log_path(invocation.prompt_path)
|
|
181
|
-
started_ts = int(time.time())
|
|
182
|
-
started_monotonic = time.monotonic()
|
|
183
|
-
status: dict[str, object] = {
|
|
184
|
-
"schemaVersion": 1,
|
|
185
|
-
"wrapper": invocation.provider.wrapper,
|
|
186
|
-
"role": invocation.role,
|
|
187
|
-
"pid": os.getpid(),
|
|
188
|
-
"started_ts": started_ts,
|
|
189
|
-
"log_path": str(log_path),
|
|
190
|
-
"stage": "started",
|
|
191
|
-
}
|
|
192
|
-
_write_status(status_path, status)
|
|
193
|
-
with log_path.open("wb") as log_file:
|
|
194
|
-
process = subprocess.Popen(
|
|
195
|
-
command,
|
|
196
|
-
cwd=invocation.execution_root,
|
|
197
|
-
stdout=subprocess.PIPE,
|
|
198
|
-
stderr=subprocess.STDOUT,
|
|
199
|
-
start_new_session=True,
|
|
200
|
-
)
|
|
201
|
-
exit_code, timed_out, idle_seconds = _stream_process(
|
|
202
|
-
process, log_file, invocation.idle_timeout_seconds
|
|
203
|
-
)
|
|
204
|
-
ended_ts = int(time.time())
|
|
205
|
-
status.update(
|
|
206
|
-
stage="exited",
|
|
207
|
-
exit_code=exit_code,
|
|
208
|
-
ended_ts=ended_ts,
|
|
209
|
-
duration_ms=int((time.monotonic() - started_monotonic) * 1000),
|
|
210
|
-
)
|
|
211
|
-
if timed_out:
|
|
212
|
-
status.update(
|
|
213
|
-
timeout=True,
|
|
214
|
-
idle_at_ts=ended_ts,
|
|
215
|
-
idle_seconds=idle_seconds,
|
|
216
|
-
terminated_by="idle-watchdog",
|
|
217
|
-
)
|
|
218
|
-
_write_status(status_path, status)
|
|
219
|
-
return exit_code
|
|
220
|
-
|
|
221
|
-
|
|
222
190
|
def main(argv: list[str]) -> int:
|
|
223
191
|
try:
|
|
224
|
-
invocation =
|
|
225
|
-
return
|
|
192
|
+
invocation = parse_invocation(argv[1:])
|
|
193
|
+
return run_worker(
|
|
194
|
+
invocation.strategy,
|
|
195
|
+
invocation.request,
|
|
196
|
+
presentation=invocation.presentation,
|
|
197
|
+
log_path=invocation.log_path,
|
|
198
|
+
status_path=invocation.status_path,
|
|
199
|
+
status_extra=invocation.status_extra,
|
|
200
|
+
)
|
|
226
201
|
except PreflightError as exc:
|
|
227
202
|
print(f"okstra-provider-exec: {exc}", file=sys.stderr)
|
|
228
203
|
return exc.exit_code
|
|
@@ -2,13 +2,20 @@
|
|
|
2
2
|
#
|
|
3
3
|
# okstra-trace-cleanup.sh — close tmux panes created during okstra runs.
|
|
4
4
|
#
|
|
5
|
-
#
|
|
6
|
-
#
|
|
7
|
-
#
|
|
8
|
-
#
|
|
9
|
-
#
|
|
10
|
-
#
|
|
11
|
-
#
|
|
5
|
+
# Worker-compute panes are tmux-pane backend siblings. Their dispatcher tags
|
|
6
|
+
# each pane it owns with a pane-level user option (`@okstra_worker_run=<RUN_DIR>`),
|
|
7
|
+
# so panes are found server-wide by tag — no tmux env var or pane-id registry is
|
|
8
|
+
# needed, and the run-scoped tag keeps concurrent okstra runs from closing each
|
|
9
|
+
# other's panes.
|
|
10
|
+
#
|
|
11
|
+
# Trace panes were `tail -F` siblings the provider wrappers split, tagged
|
|
12
|
+
# `@okstra_trace_run` / `@okstra_status`. Those wrappers are now four-line
|
|
13
|
+
# entrypoints and worker progress renders into the worker's own pane, so
|
|
14
|
+
# NOTHING SPAWNS A TRACE PANE and neither tag has a writer left. The trace
|
|
15
|
+
# paths below still run and simply match nothing; `--reclaim-completed`, which
|
|
16
|
+
# keys on `@okstra_status`, is inert for the same reason. Kept rather than
|
|
17
|
+
# deleted because the hooks that call this script are already seeded on user
|
|
18
|
+
# machines — retiring the trace machinery is a deliberate follow-up.
|
|
12
19
|
#
|
|
13
20
|
# Two invocation shapes:
|
|
14
21
|
#
|
|
@@ -1,27 +1,34 @@
|
|
|
1
1
|
#!/usr/bin/env python3
|
|
2
|
-
"""okstra-wrapper-status.py —
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
|
|
6
|
-
`
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
2
|
+
"""okstra-wrapper-status.py — standalone CLI for the worker status sidecar.
|
|
3
|
+
|
|
4
|
+
A CLI worker dispatch runs a long-running CLI in the background, and the
|
|
5
|
+
polling channel that watches it carries only stdout plus a binary
|
|
6
|
+
`running`/`completed` state. Several recovery decisions need more —
|
|
7
|
+
specifically, "did this worker start at all, when, and how did it finish?" — so
|
|
8
|
+
a small JSON sidecar is written at `<prompt-path>.status.json` that survives
|
|
9
|
+
independent of the polling channel.
|
|
10
|
+
|
|
11
|
+
**Nothing calls this script.** Every provider entrypoint now runs through
|
|
12
|
+
`scripts/okstra_ctl/worker_runner.py`, which writes the same document
|
|
13
|
+
in-process; the shell wrappers that shelled out to this file were its only
|
|
14
|
+
callers and they are gone. It is kept for now rather than deleted because
|
|
15
|
+
removing it also means touching the build sync list, the install payload, its
|
|
16
|
+
own test module and the docs that name it — a deliberate follow-up, not a side
|
|
17
|
+
effect of the wrapper migration. `scripts/okstra_ctl/wrapper_status.py` is the
|
|
18
|
+
reader, and what it reads is what the runner writes.
|
|
19
|
+
|
|
20
|
+
Consumers of the sidecar:
|
|
21
|
+
|
|
22
|
+
* A CLI worker's diagnostic step: read `log_path` to capture a tail when
|
|
23
|
+
`exit_code == 0` but the canonical Result file is absent.
|
|
24
|
+
* Lead: cross-check `started_ts` / `ended_ts` to distinguish "worker hung
|
|
18
25
|
before CLI launched" from "CLI finished but never wrote artifact" when
|
|
19
26
|
applying the redispatch policy (see team-contract "Lead Redispatch
|
|
20
27
|
Policy on Result-Missing").
|
|
21
28
|
|
|
22
|
-
Failures are deliberately non-fatal for the caller — the
|
|
23
|
-
|
|
24
|
-
|
|
29
|
+
Failures are deliberately non-fatal for the caller — the caller's main job is
|
|
30
|
+
to run the underlying CLI; a missing sidecar must not break that. On any error
|
|
31
|
+
the script prints a one-line diagnostic to stderr and exits 0.
|
|
25
32
|
|
|
26
33
|
Schema (schemaVersion 1):
|
|
27
34
|
|
|
@@ -57,7 +57,7 @@ Use the screen to tell "still working" from "stuck", and to see at a glance whic
|
|
|
57
57
|
- Worker completion is valid only from `workerDispatches[]`, terminal status sidecars, and required Result Paths. Pane creation alone is not completion.
|
|
58
58
|
- Reverify uses a fresh jobs file at `runs/<task-type>/state/reverify-jobs-r<N>-<task-type>-<seq>.json`, sets `dispatchKind: "reverify-r<N>"`, and dispatches with `okstra team dispatch --project-root <root> --run-manifest <path> --dispatch-kind reverify-r<N> --jobs-file <jobs-file>`.
|
|
59
59
|
- Report-writer uses a fresh one-job jobs file with `dispatchKind: "report-writer"` and the same schema, then dispatches through `okstra team dispatch --project-root <root> --run-manifest <path> --jobs-file <jobs-file>`.
|
|
60
|
-
- Every reverify or report-writer jobs file carries `workerId`, `provider`, `role`, `modelExecutionValue`, `promptPath`, `resultPath`, `workerResultPath`, and `completionPaths`. For reverify, set `role` to `worker-reverify-r<N>`
|
|
60
|
+
- Every reverify or report-writer jobs file carries `workerId`, `provider`, `role`, `modelExecutionValue`, `promptPath`, `resultPath`, `workerResultPath`, and `completionPaths`. For reverify, set `role` to `worker-reverify-r<N>` — the role selects the dispatch's idle budget and is recorded in the run's status sidecar, so it must name the actual assignment. The report-writer completion paths include both data.json and the worker-results audit file.
|
|
61
61
|
- After either dispatch, run `okstra team await --project-root <root> --run-manifest <path>` before evaluating terminal status or completion paths.
|
|
62
62
|
|
|
63
63
|
## Completion, cleanup, and resume
|
|
@@ -540,7 +540,7 @@ Schema rules:
|
|
|
540
540
|
|
|
541
541
|
## Coverage critic pass
|
|
542
542
|
|
|
543
|
-
Runs only when `convergence.critic.enabled == true` (set by `--critic <provider>` or the okstra-run `critic_pick` step; default off). Applies to the three finding-producing phases (`requirements-discovery`, `error-analysis`, `implementation-planning`); for `final-verification` the critic runs in a different mode — see §"Acceptance critic pass (final-verification)". This pass targets **
|
|
543
|
+
Runs only when `convergence.critic.enabled == true` (set by `--critic <provider>` or the okstra-run `critic_pick` step; default off). Applies to the three finding-producing phases (`requirements-discovery`, `error-analysis`, `implementation-planning`); for `final-verification` the critic runs in a different mode — see §"Acceptance critic pass (final-verification)". This pass targets **scope in both directions** — findings that are missing (coverage) and work the findings propose that no requirement asked for (over-scope) — distinct from convergence, which targets **agreement quality** among the findings already raised. The pass keeps its `coverage` mode id and `gaps` vocabulary for both halves; the two are told apart by each candidate's `category`, so no schema or reducer distinguishes them.
|
|
544
544
|
|
|
545
545
|
### When
|
|
546
546
|
|
|
@@ -554,10 +554,12 @@ Dispatch one fresh pass to `config.critic.provider` through `redispatch_worker`,
|
|
|
554
554
|
|
|
555
555
|
The `-worker-` token is load-bearing, not decoration: the critic prompt carries the same generated anchor headers as every other worker ([team-contract](./team-contract.md) §"Worker prompts"), and its `**Audit sidecar path:**` comes from passing that result path through `okstra_ctl.worker_artifact_paths.audit_sidecar_rel()`, which inserts `-audit-` after the token and raises without it. A `<provider>-critic-...` name leaves the lead choosing between breaking the contract and hand-inventing the sidecar name. Note that `originWorker` stays `"<provider>-critic"` — that is a worker id in the convergence state, not a filename, and the two do not have to match.
|
|
556
556
|
|
|
557
|
-
The critic prompt
|
|
557
|
+
The critic prompt carries the full analysis contract those anchor headers belong to — worker preamble, error-contract path, audit sidecar, packet boundary — but it is **exempt from the initial-analysis equality group**. Initial analysis workers must receive byte-identical normalized bodies (`worker_prompt_contract.validate_analysis_prompt_set`), and a critic body is deliberately unlike them, so `okstra_ctl.worker_prompt_policy.resolve_prompt_plan` resolves `dispatchKind = "critic"` to `equality_group = None` while keeping `audience = "analysis"`. Never reshape a critic prompt to match the initial one to satisfy that check — matching it would delete the pass. **Enforced:** `tests/contract/test_critic_prompt_equality_exemption.py` pins both directions — the critic is exempt, and two mismatched *initial* prompts still fail.
|
|
558
558
|
|
|
559
|
-
|
|
560
|
-
|
|
559
|
+
The critic prompt seeds the consolidated findings and asks for two things only — coverage gaps and unrequested work. The second half exists because every other scope check in the lifecycle runs *later*: the plan-body `P-Opt-*` verdict judges options this phase has not authored yet (report-writer writes the plan in Phase 6), and the `implementation` verifier judges a diff. Here the unit is a **finding**, which is the input those stages build on — catching unrequested work while it is still a finding is the cheapest place to catch it at all.
|
|
560
|
+
|
|
561
|
+
Required reading before proposing a gap or an over-scope candidate:
|
|
562
|
+
- the current run's `analysis-packet.md` for requirements and phase scope — for the over-scope half this is the authority you search against, so read it before judging any finding unrequested;
|
|
561
563
|
- `convergence-groups-<task-type>-<seq>.json` for the complete Round 0 ledger;
|
|
562
564
|
- every initial analysis-worker result named by team-state;
|
|
563
565
|
- each matching audit sidecar, to distinguish an uninspected path from a claim that was inspected but summarized during grouping.
|
|
@@ -565,13 +567,30 @@ Required reading before proposing a gap:
|
|
|
565
567
|
Operational guardrails are not task requirements. A gap must trace to a brief requirement, an analysis-packet scope item, a source path the packet authorizes, or an evidence claim in a worker result. Do NOT infer missing verification from a one-line summary; open the named result and audit sidecar first.
|
|
566
568
|
|
|
567
569
|
```
|
|
568
|
-
You are the
|
|
569
|
-
|
|
570
|
+
You are the scope critic for <task-key>. Below are the consolidated findings the
|
|
571
|
+
workers produced. Your job has exactly two halves. Answer both.
|
|
572
|
+
|
|
573
|
+
(1) MISSING — name what nobody covered:
|
|
570
574
|
- files / directories / execution paths nobody inspected,
|
|
571
575
|
- requirements or acceptance points with zero findings,
|
|
572
576
|
- claims raised but never verified.
|
|
573
|
-
For each
|
|
574
|
-
|
|
577
|
+
For each, emit a NEW finding with evidence (file:line or the requirement quote).
|
|
578
|
+
|
|
579
|
+
(2) UNREQUESTED — name work these findings propose that no requirement asked for:
|
|
580
|
+
- a finding whose proposed change serves no requirement, scope item, or
|
|
581
|
+
acceptance point you can QUOTE from the analysis packet,
|
|
582
|
+
- an abstraction, configuration knob, or generalization proposed for a caller or
|
|
583
|
+
a case nobody has stated,
|
|
584
|
+
- a rewrite, migration, or cleanup of code the requirements never mention.
|
|
585
|
+
For each, emit a candidate with `category: "unrequested-scope"`, quote the
|
|
586
|
+
proposed work verbatim, and state which requirement you searched for and did not
|
|
587
|
+
find.
|
|
588
|
+
|
|
589
|
+
Do NOT restate an existing finding. Judge (2) against the analysis packet's
|
|
590
|
+
requirements and scope, never against your own preference for how the code should
|
|
591
|
+
look — "I would have done it differently" is not unrequested work, and neither is
|
|
592
|
+
work the packet authorizes but you consider unnecessary. If a half has nothing,
|
|
593
|
+
say so explicitly for that half; silence on one half is an incomplete result.
|
|
575
594
|
```
|
|
576
595
|
|
|
577
596
|
### Gap verification (1 adversarial reverify round)
|
|
@@ -579,8 +598,17 @@ Each critic gap enters the verification queue as a finding with `originWorker =
|
|
|
579
598
|
|
|
580
599
|
**A gap that received no verdict is NOT a rejected gap (BLOCKING).** Dropping applies only to gaps the voters actually judged. A gap can also end the round *unjudged* — the verification dispatch returned a terminal non-result (`timeout`, `error`, no result file), the returned result covered only some of the gaps, or no non-critic analyser was available to vote at all. Nobody inspected those, so classifying them as hallucinations is a fabricated verdict. Each one MUST be recorded as a `## 5. Missing Information and Risks` row (`missingInformation`, `source: "critic-unverified"`) whose `risk` names the gap and the reason verification did not complete, and counted in `config.critic.gapsUnverified`. They are **not** promoted to findings (unverified) and **not** raised as `clarification` items — an unverified gap needs an analyser to verify it on the next run, not a decision from the user. Silently losing them is a contract violation: the batch that times out is exactly the batch of gaps too expensive to check, so the highest-risk items are the ones that vanish.
|
|
581
600
|
|
|
601
|
+
**`category: "unrequested-scope"` candidates are classified the same way but disposed of differently.** A coverage gap the voters contest is a hallucination — nothing was actually missing, so dropping it costs one wasted verification. An over-scope candidate the voters contest is a *disagreement about whether the work was asked for*, and dropping that silently returns the run to the state this half exists to change. So:
|
|
602
|
+
|
|
603
|
+
- `full-consensus` / `partial-consensus` → merge as a finding, exactly like a coverage gap. The merged finding names the unrequested work and the requirement search that came up empty; the phase's own deliverable rules decide whether it lands as a dropped item or a clarification row.
|
|
604
|
+
- `contested` / `worker-unique` → counted in `config.critic.gapsRejected` (unchanged accounting) but ALSO recorded as a `## 5. Missing Information and Risks` row with `source: "critic-unconfirmed-scope"`, whose `risk` quotes the proposed work and the split verdict. It is **not** promoted to a finding: a contested over-scope claim must not block a plan on one worker's taste. The user reads the row and decides.
|
|
605
|
+
- no verdict at all → the `critic-unverified` rule above applies unchanged.
|
|
606
|
+
|
|
607
|
+
The asymmetry is deliberate and runs the opposite way from the coverage half: a false "you missed something" costs a verification, while a false "you built too much" costs a real requirement — so the first may be dropped outright and the second is recorded either way. This is the same reasoning that makes the `final-verification` acceptance critic never drop a candidate, applied to the one direction where dropping is otherwise the default.
|
|
608
|
+
|
|
582
609
|
### State
|
|
583
610
|
- `convergence.critic` manifest block: `{ enabled, provider, modelExecutionValue }`.
|
|
611
|
+
- Each candidate's `category` tells the two halves apart: literal `"unrequested-scope"` for the over-scope half, any other value for a coverage gap. `schemas/convergence-critic-results-v1.0.schema.json` leaves `category` a free string, so this needs no schema or reducer change — but it also means nothing machine-checks the spelling. A misspelled category is read as a coverage gap and silently takes the drop-on-contested path.
|
|
584
612
|
- The lead passes one canonical coverage batch with `{ schemaVersion, taskKey, mode, provider, modelExecutionValue, dispatches, gaps }`; each gap carries its candidate fields plus `gapId` and `votes`. `dispatches[]` contains exactly one row for every Phase 4 analyser except the critic, even when execution did not produce a result: persist `status: timeout | error | not-run` and the elapsed `durationMs` instead of omitting that analyser. `apply-critic-gaps` rejects a non-terminal main queue, duplicate or missing analysers, unknown workers, critic dispatches/votes, votes without a completed dispatch, and a second batch.
|
|
585
613
|
- Convergence state artifact: merged gaps appear in `findings[]` with `source: "critic"` and `rounds: []`. The separate `criticVerification.gaps[]` ledger retains each gap's `summary`, `category`, `ticketIds`, `originEvidence`, optional `evidenceArtifacts`, classification, merge link, and votes. Strict v1.3 validation deterministically replays each complete critic-origin finding from that ledger; a critic batch never increments `roundHistory` or `totalRounds` and never creates a fake main round.
|
|
586
614
|
- `config.critic` is `{ provider, modelExecutionValue, gapsProposed, gapsMerged, gapsRejected, gapsUnverified }`, with `gapsProposed = gapsMerged + gapsRejected + gapsUnverified`. `full-consensus` / `partial-consensus` gaps merge, `contested` / `worker-unique` gaps count as rejected, and gaps with no usable analyser vote appear in both the ledger and final `unverifiedGaps[]`.
|
|
@@ -104,9 +104,10 @@ Required checkpoints:
|
|
|
104
104
|
- `PROGRESS: phase-5-collect worker=<role> status=<terminal-status>` — once per worker, immediately after the result file is verified.
|
|
105
105
|
- `PROGRESS: phase-5.5-convergence round=<N> queue=<count>` — at the start of each convergence round (Phase 5.5).
|
|
106
106
|
- `PROGRESS: phase-5.6-critic provider=<provider> gaps=<n>` — after the critic result is collected (Phase 5.6, opt-in; the critic dispatch itself fires concurrently with the first 5.5 reverify round). Omitted when `convergence.critic.enabled == false`.
|
|
107
|
-
- `PROGRESS: phase-batch-cleanup panes=<n>` — immediately after cleaning up the previous batch's panes, at each batch boundary (① just before the first `phase-5.5-convergence` round ② just before the `phase-6-synthesis` report-writer dispatch). `<n>` is the number of panes reclaimed at that boundary —
|
|
107
|
+
- `PROGRESS: phase-batch-cleanup panes=<n>` — immediately after cleaning up the previous batch's panes, at each batch boundary (① just before the first `phase-5.5-convergence` round ② just before the `phase-6-synthesis` report-writer dispatch). `<n>` is the number of panes reclaimed at that boundary — worker-compute panes plus completed teammate panes, which are panes too — read from the cleanup's `--list` pass taken immediately before the reclaim, never estimated. Expose only the counts and NEVER expose `%NNN`/lead-pane.id/raw worker handles. Just before the first batch (analysis-worker dispatch) there is nothing to clean up, so it is a no-op and the marker is omitted.
|
|
108
108
|
- `PROGRESS: phase-6-synthesis dispatching report-writer-worker` — at the start of Phase 6.
|
|
109
109
|
- `PROGRESS: phase-5.5.9-plan-verify round=<N> items=<count>` — immediately before dispatching each plan-body verification round (`implementation-planning` only; see [plan-body-verification](./plan-body-verification.md) §"Round protocol"). Each round is a worker batch like any other, so round 2 and later MUST be preceded by a `phase-batch-cleanup` line reclaiming the previous round's verifiers. The numbering keeps this line sorted where the work happens — after Phase 6, because the round verifies the drafted plan body.
|
|
110
|
+
- `PROGRESS: user-confirm <C-NNN> <the question, one line>` — immediately before asking the user about anything that would otherwise become an open `Blocks=approval` row (see "User confirmation before an approval blocker" below). Not tied to a phase: it fires wherever the blocker surfaces. `<C-NNN>` is the id the row will carry, so the answer and the row can be matched afterwards.
|
|
110
111
|
- `PROGRESS: phase-7-persist updating manifests` — at the start of Phase 7.
|
|
111
112
|
- `PROGRESS: phase-7-teardown shutting-down-workers` — only after usage collection and user approval, immediately before `shutdown_workers`; omitted when no cleanup resource exists or the user keeps it.
|
|
112
113
|
- `PROGRESS: complete final-report=<relative-path>` — final summary line, after all persistence.
|
|
@@ -117,6 +118,27 @@ Do NOT replace them with prose ("Now I'm starting Phase 2..."), do NOT skip a ch
|
|
|
117
118
|
|
|
118
119
|
**Enforcement:** the Phase 7 validator (`validators/validate-run.py` → `validate_session_conformance.py`) reads the selected adapter's declared conformance evidence/event source within the run window and fails the run as `contract-violated` when a required checkpoint is missing — including the per-worker `phase-4-dispatch` / `phase-5-collect` lines (which must name each worker's role) and the `phase-batch-cleanup` lines that MUST precede the first `phase-5.5-convergence` round and the `phase-6-synthesis` report-writer dispatch. When the plan-body state file records two or more rounds, `_check_plan_verify_cleanup_checkpoints` additionally requires a `phase-5.5.9-plan-verify` line per round and a `phase-batch-cleanup` between consecutive rounds — that boundary sat outside both older checks, so a five-round self-fix loop left every round's verifiers holding their panes. `phase-7-teardown` and `complete` fire after validation and are not checked.
|
|
119
120
|
|
|
121
|
+
## User confirmation before an approval blocker (BLOCKING)
|
|
122
|
+
|
|
123
|
+
An open `Blocks=approval` row stops the whole task: the plan cannot be approved, `implementation` cannot start, and the answer arrives only through a separate user-response cycle and a re-run. It is the most expensive artifact this contract lets the lead produce. **Before writing one, ask the user.**
|
|
124
|
+
|
|
125
|
+
This is not a phase. It fires wherever the blocker surfaces — during intake when the directive and the brief disagree, mid-convergence when workers split on something only the user can settle, in the §5.5.9 self-fix loop when an item no round can clear keeps the gate red.
|
|
126
|
+
|
|
127
|
+
The sequence is fixed:
|
|
128
|
+
|
|
129
|
+
1. Emit `PROGRESS: user-confirm <C-NNN> <the question, one line>` with the id the row would carry.
|
|
130
|
+
2. Ask in plain user-facing text: what is undecided, the options with their consequences, and which one you recommend. One question at a time.
|
|
131
|
+
3. On an answer — record it in the row's `userInput`, set `status: answered` and `userConfirmation: asked-and-answered`, apply it, and **keep going in this run**. An answered question is not a blocker, and a run that stops anyway wastes the answer it just received.
|
|
132
|
+
4. Only when asking fails does the row stay open: `asked-awaiting` when the user has not answered, `deferred-no-interactive-session` when this run has no user to ask.
|
|
133
|
+
|
|
134
|
+
**Predicting the blocker is not the same as raising it.** A lead that says "this will likely become an approval blocker; I will ask at that point" has already reached the moment — ask then, in that message. One run announced exactly that, never asked, wrote the row anyway, and then spent its entire self-fix budget on a gate no round could clear, because the user had already answered the question before the run started.
|
|
135
|
+
|
|
136
|
+
**`lead-directed` blockers cannot be deferred.** When the item is the lead's own judgment rather than a worker's finding, and this run has nobody to ask, the row is not the outlet — record a Working Assumption in `## 5. Missing Information and Risks` naming the assumption the plan proceeds under, exactly as a surviving planner-fixable item does, and let the plan proceed. Blocking a plan on the lead's own judgment in a run where that judgment cannot be put to the user only moves the work to a re-run.
|
|
137
|
+
|
|
138
|
+
**Never seed the answer into a worker prompt.** Instructing a worker to "raise this as a user decision rather than choosing" and then reporting the resulting agreement as an independent finding misrepresents where the blocker came from. If it is the lead's judgment, `origin` is `lead-directed` — see [_common-contract.md](../profiles/_common-contract.md) "Clarification request policy".
|
|
139
|
+
|
|
140
|
+
**Enforcement:** `validators/validate-run.py` `_validate_open_approval_blocker_provenance` fails any open approval blocker missing `origin` / `userConfirmation`, and any `lead-directed` one deferred for want of an interactive session; `validators/validate_session_conformance.py` `_check_user_confirm_checkpoints` fails a row claiming the user was asked when no matching `user-confirm` line exists in this run's evidence.
|
|
141
|
+
|
|
120
142
|
## Model assignments
|
|
121
143
|
|
|
122
144
|
**The lead never invents a model.** Every role's model is read from `task-manifest.json` → `resultContract.requiredWorkerRoles[*].modelExecutionValue` (and the lead model metadata). A missing assignment is a manifest defect, not a license to fall back — see [team-contract](./team-contract.md) "Model Assignment Rules". The manifest is always populated at run-prep time by the CLI, which seeds these values from `OKSTRA_DEFAULT_*_MODEL` (`scripts/okstra_ctl/run.py`).
|