okstra 0.164.0 → 0.165.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +1 -1
- package/docs/architecture.md +12 -8
- package/docs/cli.md +7 -3
- package/docs/for-ai/README.md +2 -2
- package/docs/for-ai/skills/okstra-inspect.md +2 -2
- package/docs/for-ai/skills/okstra-user-response.md +2 -2
- package/docs/project-structure-overview.md +15 -9
- package/package.json +1 -1
- package/runtime/BUILD.json +2 -2
- package/runtime/agents/workers/antigravity-worker.md +9 -7
- package/runtime/agents/workers/codex-worker.md +9 -7
- package/runtime/agents/workers/grok-worker.md +6 -4
- package/runtime/agents/workers/kimi-worker.md +6 -4
- package/runtime/bin/okstra-antigravity-exec.sh +1 -340
- package/runtime/bin/okstra-claude-exec.sh +1 -178
- package/runtime/bin/okstra-codex-exec.sh +1 -467
- package/runtime/bin/okstra-provider-exec.py +165 -190
- package/runtime/bin/okstra-spawn-followups.py +29 -1
- package/runtime/bin/okstra-trace-cleanup.sh +14 -7
- package/runtime/bin/okstra-wrapper-status.py +26 -19
- package/runtime/prompts/lead/adapters/cmux.md +1 -1
- package/runtime/prompts/lead/convergence.md +36 -8
- package/runtime/prompts/lead/okstra-lead-contract.md +23 -1
- package/runtime/prompts/lead/plan-body-verification.md +9 -1
- package/runtime/prompts/lead/report-writer.md +1 -0
- package/runtime/prompts/lead/team-contract.md +3 -3
- package/runtime/prompts/profiles/_common-contract.md +9 -1
- package/runtime/prompts/profiles/_coverage-critic.md +1 -1
- package/runtime/prompts/profiles/_implementation-diff-review.md +3 -1
- package/runtime/prompts/profiles/_implementation-self-check.md +1 -1
- package/runtime/prompts/profiles/_implementation-verifier.md +3 -1
- package/runtime/prompts/profiles/implementation-planning.md +5 -3
- package/runtime/python/okstra_ctl/adapters/hosts/claude-code/relay.md +1 -1
- package/runtime/python/okstra_ctl/adapters/hosts/external/relay.md +1 -1
- package/runtime/python/okstra_ctl/adapters/providers/antigravity/adapter.py +148 -0
- package/runtime/python/okstra_ctl/adapters/providers/claude/adapter.py +55 -0
- package/runtime/python/okstra_ctl/adapters/providers/codex/adapter.py +41 -0
- package/runtime/python/okstra_ctl/adapters/providers/grok/adapter.py +44 -0
- package/runtime/python/okstra_ctl/adapters/providers/kimi/adapter.py +42 -0
- package/runtime/python/okstra_ctl/dispatch_core.py +5 -1
- package/runtime/python/okstra_ctl/dispatch_state.py +10 -0
- package/runtime/python/okstra_ctl/domain/provider.py +5 -1
- package/runtime/python/okstra_ctl/domain/worker_exec.py +102 -0
- package/runtime/python/okstra_ctl/domain/worker_role.py +34 -0
- package/runtime/python/okstra_ctl/domain/worker_stream.py +261 -0
- package/runtime/python/okstra_ctl/incremental_scope.py +16 -4
- package/runtime/python/okstra_ctl/report_html/common.py +71 -25
- package/runtime/python/okstra_ctl/report_html/models.py +5 -0
- package/runtime/python/okstra_ctl/report_html/render.py +1 -1
- package/runtime/python/okstra_ctl/report_html/run_usage.py +19 -0
- package/runtime/python/okstra_ctl/report_html/view_models/implementation_planning.py +14 -0
- package/runtime/python/okstra_ctl/report_views.py +44 -16
- package/runtime/python/okstra_ctl/stage_citations.py +52 -15
- package/runtime/python/okstra_ctl/user_response.py +45 -29
- package/runtime/python/okstra_ctl/wizard.py +23 -13
- package/runtime/python/okstra_ctl/worker_prompt_policy.py +10 -3
- package/runtime/python/okstra_ctl/worker_request.py +140 -0
- package/runtime/python/okstra_ctl/worker_runner.py +622 -0
- package/runtime/python/okstra_token_usage/collect.py +8 -1
- package/runtime/python/okstra_token_usage/report.py +42 -0
- package/runtime/python/okstra_token_usage/task_totals.py +88 -0
- package/runtime/schemas/final-report-v1.0.schema.json +70 -0
- package/runtime/schemas/final-report-v2.0.schema.json +90 -0
- package/runtime/skills/okstra-inspect/SKILL.md +1 -2
- package/runtime/skills/okstra-inspect/facets/logs.md +5 -5
- package/runtime/skills/okstra-inspect/facets/run-audit.md +3 -3
- package/runtime/skills/okstra-run/SKILL.md +1 -1
- package/runtime/skills/okstra-user-response/SKILL.md +15 -5
- package/runtime/templates/report-writer-prompt-preamble.md +1 -0
- package/runtime/templates/reports/html/assets/base.css +8 -4
- package/runtime/templates/reports/html/base.template.html +12 -6
- package/runtime/templates/reports/html/i18n/en.json +29 -6
- package/runtime/templates/reports/html/i18n/ko.json +29 -6
- package/runtime/templates/reports/html/macros/forms.html +9 -3
- package/runtime/templates/reports/html/tasks/implementation-planning.template.html +14 -19
- package/runtime/templates/reports/report.js +59 -26
- package/runtime/templates/reports/user-response.template.md +12 -8
- package/runtime/validators/validate-run.py +88 -7
- package/runtime/validators/validate_session_conformance.py +62 -1
- package/src/cli-registry.mjs +0 -7
- package/runtime/bin/okstra-wrapper-agy-stream.py +0 -61
- package/runtime/python/okstra_ctl/error_issue.py +0 -640
- package/runtime/python/okstra_ctl/issue_signals.py +0 -186
- package/runtime/skills/okstra-inspect/facets/error-issue.md +0 -77
- package/src/commands/inspect/error-issue.mjs +0 -27
|
@@ -1,121 +1,184 @@
|
|
|
1
1
|
#!/usr/bin/env python3
|
|
2
|
-
"""Run an external LLM CLI with the shared okstra wrapper contract.
|
|
2
|
+
"""Run an external LLM CLI with the shared okstra wrapper contract.
|
|
3
|
+
|
|
4
|
+
Argument parsing and provider lookup only — the run itself belongs to
|
|
5
|
+
``okstra_ctl.worker_runner``, which every provider shares. What stays here is
|
|
6
|
+
the request the provider strategies then take on trust: resolved paths, the
|
|
7
|
+
write scope in the order the CLIs are told it, and the role's idle budget.
|
|
8
|
+
"""
|
|
3
9
|
from __future__ import annotations
|
|
4
10
|
|
|
5
|
-
import json
|
|
6
11
|
import os
|
|
7
|
-
import selectors
|
|
8
12
|
import shutil
|
|
9
|
-
import signal
|
|
10
|
-
import subprocess
|
|
11
13
|
import sys
|
|
12
|
-
import time
|
|
13
14
|
from dataclasses import dataclass
|
|
14
15
|
from pathlib import Path
|
|
15
|
-
from typing import
|
|
16
|
+
from typing import Any, Mapping
|
|
17
|
+
|
|
18
|
+
_HERE = Path(__file__).resolve().parent
|
|
19
|
+
# ``okstra_ctl`` sits beside this file in the repo (``scripts/``) but under
|
|
20
|
+
# ``~/.okstra/lib/python/`` once installed, and the four-line shell entrypoints
|
|
21
|
+
# that exec this script set no PYTHONPATH. Offer both, repo first.
|
|
22
|
+
_HOME_LIB = (
|
|
23
|
+
Path(os.environ.get("OKSTRA_HOME", str(Path.home() / ".okstra"))) / "lib" / "python"
|
|
24
|
+
)
|
|
25
|
+
sys.path.insert(0, str(_HERE))
|
|
26
|
+
if _HOME_LIB.is_dir() and str(_HOME_LIB) not in sys.path:
|
|
27
|
+
sys.path.append(str(_HOME_LIB))
|
|
28
|
+
|
|
29
|
+
from okstra_ctl.domain.provider import ProviderSpec, UnknownProviderError # noqa: E402
|
|
30
|
+
from okstra_ctl.domain.worker_exec import ( # noqa: E402
|
|
31
|
+
ExecutionStrategy,
|
|
32
|
+
WorkerExecRequest,
|
|
33
|
+
)
|
|
34
|
+
from okstra_ctl.registry.provider_registry import ( # noqa: E402
|
|
35
|
+
default_provider_registry,
|
|
36
|
+
)
|
|
37
|
+
from okstra_ctl.worker_request import build_request, idle_timeout # noqa: E402
|
|
38
|
+
from okstra_ctl.worker_runner import LIVE, QUIET, run_worker # noqa: E402
|
|
39
|
+
|
|
40
|
+
_USAGE = (
|
|
41
|
+
"usage: okstra-provider-exec.py <provider> <project-root> "
|
|
42
|
+
"<model-execution-value> <prompt-path> [worktree-path] [role] "
|
|
43
|
+
"[idle-timeout-seconds] [--presentation live|quiet]"
|
|
44
|
+
)
|
|
45
|
+
|
|
46
|
+
_PRESENTATION_FLAG = "--presentation"
|
|
47
|
+
_PRESENTATIONS = (LIVE, QUIET)
|
|
48
|
+
|
|
49
|
+
|
|
50
|
+
class PreflightError(Exception):
|
|
51
|
+
def __init__(self, exit_code: int, message: str) -> None:
|
|
52
|
+
super().__init__(message)
|
|
53
|
+
self.exit_code = exit_code
|
|
16
54
|
|
|
17
55
|
|
|
18
56
|
@dataclass(frozen=True)
|
|
19
|
-
class
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
57
|
+
class Invocation:
|
|
58
|
+
strategy: ExecutionStrategy
|
|
59
|
+
request: WorkerExecRequest
|
|
60
|
+
presentation: str
|
|
61
|
+
log_path: Path
|
|
62
|
+
status_path: Path
|
|
63
|
+
status_extra: Mapping[str, Any]
|
|
64
|
+
|
|
65
|
+
|
|
66
|
+
def parse_invocation(argv: list[str]) -> Invocation:
|
|
67
|
+
"""Resolve the wrapper's positional contract into one runnable dispatch."""
|
|
68
|
+
positional, presentation = _take_presentation(argv)
|
|
69
|
+
if not 4 <= len(positional) <= 7:
|
|
70
|
+
raise PreflightError(64, _USAGE)
|
|
71
|
+
provider_id, project_root_raw, model, prompt_raw = positional[:4]
|
|
72
|
+
worktree_raw = positional[4] if len(positional) >= 5 else ""
|
|
73
|
+
role = positional[5] if len(positional) >= 6 and positional[5] else "worker"
|
|
74
|
+
timeout_raw = positional[6] if len(positional) >= 7 and positional[6] else ""
|
|
75
|
+
|
|
76
|
+
spec = _provider_spec(provider_id)
|
|
77
|
+
project_root = _existing_dir(project_root_raw, 65, "project-root")
|
|
78
|
+
if not model:
|
|
79
|
+
raise PreflightError(66, "model-execution-value is empty")
|
|
80
|
+
prompt_path = _existing_file(prompt_raw, 67, "prompt-path")
|
|
81
|
+
idle_timeout_seconds = _idle_timeout(timeout_raw, role)
|
|
82
|
+
worktree = (
|
|
83
|
+
_existing_dir(worktree_raw, 68, "worktree-path") if worktree_raw else None
|
|
84
|
+
)
|
|
23
85
|
|
|
86
|
+
request = build_request(
|
|
87
|
+
prompt_text=prompt_path.read_text(encoding="utf-8"),
|
|
88
|
+
model=model,
|
|
89
|
+
project_root=project_root,
|
|
90
|
+
worktree_path=worktree,
|
|
91
|
+
role=role,
|
|
92
|
+
idle_timeout_seconds=idle_timeout_seconds,
|
|
93
|
+
)
|
|
94
|
+
strategy = spec.exec_strategy
|
|
95
|
+
_check_command(strategy, request)
|
|
96
|
+
return Invocation(
|
|
97
|
+
strategy=strategy,
|
|
98
|
+
request=request,
|
|
99
|
+
presentation=presentation,
|
|
100
|
+
log_path=_log_path(prompt_path),
|
|
101
|
+
status_path=Path(f"{prompt_path}.status.json"),
|
|
102
|
+
status_extra={"wrapper": spec.wrapper, "role": role},
|
|
103
|
+
)
|
|
24
104
|
|
|
25
|
-
def _grok_args(prompt: str, model: str, cwd: str) -> list[str]:
|
|
26
|
-
return [
|
|
27
|
-
"grok",
|
|
28
|
-
"-p",
|
|
29
|
-
prompt,
|
|
30
|
-
"-m",
|
|
31
|
-
model,
|
|
32
|
-
"--output-format",
|
|
33
|
-
"streaming-json",
|
|
34
|
-
"--cwd",
|
|
35
|
-
cwd,
|
|
36
|
-
]
|
|
37
105
|
|
|
106
|
+
def _take_presentation(argv: list[str]) -> tuple[list[str], str]:
|
|
107
|
+
"""Split the one flag out of an otherwise positional argv.
|
|
108
|
+
|
|
109
|
+
Defaults to ``quiet``. ``live`` is only ever right where a screen was
|
|
110
|
+
declared, and the only callers that can declare one are the pane backends —
|
|
111
|
+
which pass the flag explicitly. Defaulting the other way assumed a screen
|
|
112
|
+
that a subagent dispatch does not have, and sent every worker's progress
|
|
113
|
+
into its caller's context window instead.
|
|
114
|
+
"""
|
|
115
|
+
positional: list[str] = []
|
|
116
|
+
presentation = QUIET
|
|
117
|
+
index = 0
|
|
118
|
+
while index < len(argv):
|
|
119
|
+
if argv[index] != _PRESENTATION_FLAG:
|
|
120
|
+
positional.append(argv[index])
|
|
121
|
+
index += 1
|
|
122
|
+
continue
|
|
123
|
+
if index + 1 >= len(argv):
|
|
124
|
+
raise PreflightError(64, f"{_PRESENTATION_FLAG} needs a value: {_USAGE}")
|
|
125
|
+
presentation = argv[index + 1]
|
|
126
|
+
index += 2
|
|
127
|
+
if presentation not in _PRESENTATIONS:
|
|
128
|
+
allowed = " | ".join(_PRESENTATIONS)
|
|
129
|
+
raise PreflightError(
|
|
130
|
+
64, f"unsupported presentation {presentation!r}. Allowed values: {allowed}"
|
|
131
|
+
)
|
|
132
|
+
return positional, presentation
|
|
38
133
|
|
|
39
|
-
def _kimi_args(prompt: str, model: str, _cwd: str) -> list[str]:
|
|
40
|
-
return ["kimi", "-p", prompt, "-m", model, "--output-format", "stream-json"]
|
|
41
134
|
|
|
135
|
+
def _provider_spec(provider_id: str) -> ProviderSpec:
|
|
136
|
+
try:
|
|
137
|
+
spec = default_provider_registry().resolve(provider_id)
|
|
138
|
+
except UnknownProviderError as exc:
|
|
139
|
+
raise PreflightError(64, str(exc)) from exc
|
|
140
|
+
if spec.exec_strategy is None:
|
|
141
|
+
raise PreflightError(
|
|
142
|
+
64, f"provider {provider_id!r} has no execution strategy to run"
|
|
143
|
+
)
|
|
144
|
+
return spec
|
|
42
145
|
|
|
43
|
-
PROVIDERS = {
|
|
44
|
-
"grok": ProviderCommand("grok", "okstra-grok-exec.sh", _grok_args),
|
|
45
|
-
"kimi": ProviderCommand("kimi", "okstra-kimi-exec.sh", _kimi_args),
|
|
46
|
-
}
|
|
47
146
|
|
|
147
|
+
def _existing_dir(raw: str, exit_code: int, label: str) -> Path:
|
|
148
|
+
"""Existence only — `build_request` owns the resolving."""
|
|
149
|
+
path = Path(raw) if raw else None
|
|
150
|
+
if path is None or not path.is_dir():
|
|
151
|
+
raise PreflightError(
|
|
152
|
+
exit_code, f"{label} is missing or not a directory: {raw!r}"
|
|
153
|
+
)
|
|
154
|
+
return path
|
|
48
155
|
|
|
49
|
-
@dataclass(frozen=True)
|
|
50
|
-
class Invocation:
|
|
51
|
-
provider: ProviderCommand
|
|
52
|
-
project_root: Path
|
|
53
|
-
model: str
|
|
54
|
-
prompt_path: Path
|
|
55
|
-
execution_root: Path
|
|
56
|
-
role: str
|
|
57
|
-
idle_timeout_seconds: int
|
|
58
156
|
|
|
157
|
+
def _existing_file(raw: str, exit_code: int, label: str) -> Path:
|
|
158
|
+
path = Path(raw) if raw else None
|
|
159
|
+
if path is None or not path.is_file():
|
|
160
|
+
raise PreflightError(exit_code, f"{label} is missing or not a file: {raw!r}")
|
|
161
|
+
return path.resolve()
|
|
59
162
|
|
|
60
|
-
class PreflightError(Exception):
|
|
61
|
-
def __init__(self, exit_code: int, message: str) -> None:
|
|
62
|
-
super().__init__(message)
|
|
63
|
-
self.exit_code = exit_code
|
|
64
163
|
|
|
164
|
+
def _idle_timeout(raw: str, role: str) -> int:
|
|
165
|
+
try:
|
|
166
|
+
return idle_timeout(raw, role)
|
|
167
|
+
except ValueError as exc:
|
|
168
|
+
raise PreflightError(69, str(exc)) from exc
|
|
65
169
|
|
|
66
|
-
def _parse_invocation(argv: list[str]) -> Invocation:
|
|
67
|
-
if len(argv) < 4 or len(argv) > 7:
|
|
68
|
-
raise PreflightError(
|
|
69
|
-
64,
|
|
70
|
-
"usage: okstra-provider-exec.py <provider> <project-root> <model-execution-value> "
|
|
71
|
-
"<prompt-path> [worktree-path] [role] [idle-timeout-seconds]",
|
|
72
|
-
)
|
|
73
|
-
provider_id, project_root_raw, model, prompt_raw = argv[:4]
|
|
74
|
-
provider = PROVIDERS.get(provider_id)
|
|
75
|
-
if provider is None:
|
|
76
|
-
raise PreflightError(64, f"unsupported provider: {provider_id}")
|
|
77
|
-
worktree_raw = argv[4] if len(argv) >= 5 else ""
|
|
78
|
-
role = argv[5] if len(argv) >= 6 and argv[5] else "worker"
|
|
79
|
-
default_timeout = 1500 if role in {"executor", "verifier"} else 600
|
|
80
|
-
timeout_raw = argv[6] if len(argv) >= 7 else str(default_timeout)
|
|
81
|
-
return _validate_invocation(
|
|
82
|
-
provider, project_root_raw, model, prompt_raw, worktree_raw, role, timeout_raw
|
|
83
|
-
)
|
|
84
170
|
|
|
171
|
+
def _check_command(strategy: ExecutionStrategy, request: WorkerExecRequest) -> None:
|
|
172
|
+
"""Refuse a missing CLI before the run leaves any artifact behind.
|
|
85
173
|
|
|
86
|
-
|
|
87
|
-
provider
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
) -> Invocation:
|
|
95
|
-
project_root = Path(project_root_raw)
|
|
96
|
-
prompt_path = Path(prompt_raw)
|
|
97
|
-
if not project_root_raw or not project_root.is_dir():
|
|
98
|
-
raise PreflightError(65, f"project-root is missing or not a directory: {project_root_raw!r}")
|
|
99
|
-
if not model:
|
|
100
|
-
raise PreflightError(66, "model-execution-value is empty")
|
|
101
|
-
if not prompt_raw or not prompt_path.is_file():
|
|
102
|
-
raise PreflightError(67, f"prompt-path is missing or not a file: {prompt_raw!r}")
|
|
103
|
-
if not timeout_raw.isdigit():
|
|
104
|
-
raise PreflightError(69, f"idle-timeout-seconds must be a non-negative integer: {timeout_raw!r}")
|
|
105
|
-
execution_root = Path(worktree_raw) if worktree_raw else project_root
|
|
106
|
-
if worktree_raw and not execution_root.is_dir():
|
|
107
|
-
raise PreflightError(68, f"worktree-path was provided but is not a directory: {worktree_raw!r}")
|
|
108
|
-
if shutil.which(provider.binary) is None:
|
|
109
|
-
raise PreflightError(127, f"{provider.binary} CLI is not installed on PATH")
|
|
110
|
-
return Invocation(
|
|
111
|
-
provider=provider,
|
|
112
|
-
project_root=project_root.resolve(),
|
|
113
|
-
model=model,
|
|
114
|
-
prompt_path=prompt_path.resolve(),
|
|
115
|
-
execution_root=execution_root.resolve(),
|
|
116
|
-
role=role,
|
|
117
|
-
idle_timeout_seconds=int(timeout_raw),
|
|
118
|
-
)
|
|
174
|
+
Building the command is the only truthful way to learn which binary this
|
|
175
|
+
provider runs, and it is a pure call. The check stays out of the runner so a
|
|
176
|
+
refused dispatch leaves no `started` status sidecar for the liveness probe to
|
|
177
|
+
read as a worker that launched.
|
|
178
|
+
"""
|
|
179
|
+
binary = strategy.build_command(request).argv[0]
|
|
180
|
+
if shutil.which(binary) is None:
|
|
181
|
+
raise PreflightError(127, f"{binary} CLI is not installed on PATH")
|
|
119
182
|
|
|
120
183
|
|
|
121
184
|
def _log_path(prompt_path: Path) -> Path:
|
|
@@ -124,105 +187,17 @@ def _log_path(prompt_path: Path) -> Path:
|
|
|
124
187
|
return Path(f"{prompt_path}.log")
|
|
125
188
|
|
|
126
189
|
|
|
127
|
-
def _write_status(path: Path, status: dict[str, object]) -> None:
|
|
128
|
-
temporary = Path(f"{path}.tmp")
|
|
129
|
-
temporary.write_text(json.dumps(status, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
|
130
|
-
os.replace(temporary, path)
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
def _terminate_process(process: subprocess.Popen[bytes]) -> None:
|
|
134
|
-
try:
|
|
135
|
-
os.killpg(process.pid, signal.SIGTERM)
|
|
136
|
-
except ProcessLookupError:
|
|
137
|
-
return
|
|
138
|
-
try:
|
|
139
|
-
process.wait(timeout=5)
|
|
140
|
-
except subprocess.TimeoutExpired:
|
|
141
|
-
try:
|
|
142
|
-
os.killpg(process.pid, signal.SIGKILL)
|
|
143
|
-
except ProcessLookupError:
|
|
144
|
-
pass
|
|
145
|
-
process.wait()
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
def _stream_process(
|
|
149
|
-
process: subprocess.Popen[bytes], log_file, idle_timeout_seconds: int
|
|
150
|
-
) -> tuple[int, bool, int]:
|
|
151
|
-
selector = selectors.DefaultSelector()
|
|
152
|
-
assert process.stdout is not None
|
|
153
|
-
selector.register(process.stdout, selectors.EVENT_READ)
|
|
154
|
-
last_output = time.monotonic()
|
|
155
|
-
timed_out = False
|
|
156
|
-
idle_seconds = 0
|
|
157
|
-
while selector.get_map():
|
|
158
|
-
for key, _ in selector.select(timeout=0.25):
|
|
159
|
-
chunk = os.read(key.fd, 8192)
|
|
160
|
-
if not chunk:
|
|
161
|
-
selector.unregister(key.fileobj)
|
|
162
|
-
continue
|
|
163
|
-
last_output = time.monotonic()
|
|
164
|
-
sys.stdout.buffer.write(chunk)
|
|
165
|
-
sys.stdout.buffer.flush()
|
|
166
|
-
log_file.write(chunk)
|
|
167
|
-
log_file.flush()
|
|
168
|
-
idle_seconds = int(time.monotonic() - last_output)
|
|
169
|
-
if idle_timeout_seconds and idle_seconds >= idle_timeout_seconds and process.poll() is None:
|
|
170
|
-
timed_out = True
|
|
171
|
-
_terminate_process(process)
|
|
172
|
-
exit_code = process.wait()
|
|
173
|
-
return (124 if timed_out else exit_code), timed_out, idle_seconds
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
def _run(invocation: Invocation) -> int:
|
|
177
|
-
prompt = invocation.prompt_path.read_text(encoding="utf-8")
|
|
178
|
-
command = invocation.provider.build_args(prompt, invocation.model, str(invocation.execution_root))
|
|
179
|
-
status_path = Path(f"{invocation.prompt_path}.status.json")
|
|
180
|
-
log_path = _log_path(invocation.prompt_path)
|
|
181
|
-
started_ts = int(time.time())
|
|
182
|
-
started_monotonic = time.monotonic()
|
|
183
|
-
status: dict[str, object] = {
|
|
184
|
-
"schemaVersion": 1,
|
|
185
|
-
"wrapper": invocation.provider.wrapper,
|
|
186
|
-
"role": invocation.role,
|
|
187
|
-
"pid": os.getpid(),
|
|
188
|
-
"started_ts": started_ts,
|
|
189
|
-
"log_path": str(log_path),
|
|
190
|
-
"stage": "started",
|
|
191
|
-
}
|
|
192
|
-
_write_status(status_path, status)
|
|
193
|
-
with log_path.open("wb") as log_file:
|
|
194
|
-
process = subprocess.Popen(
|
|
195
|
-
command,
|
|
196
|
-
cwd=invocation.execution_root,
|
|
197
|
-
stdout=subprocess.PIPE,
|
|
198
|
-
stderr=subprocess.STDOUT,
|
|
199
|
-
start_new_session=True,
|
|
200
|
-
)
|
|
201
|
-
exit_code, timed_out, idle_seconds = _stream_process(
|
|
202
|
-
process, log_file, invocation.idle_timeout_seconds
|
|
203
|
-
)
|
|
204
|
-
ended_ts = int(time.time())
|
|
205
|
-
status.update(
|
|
206
|
-
stage="exited",
|
|
207
|
-
exit_code=exit_code,
|
|
208
|
-
ended_ts=ended_ts,
|
|
209
|
-
duration_ms=int((time.monotonic() - started_monotonic) * 1000),
|
|
210
|
-
)
|
|
211
|
-
if timed_out:
|
|
212
|
-
status.update(
|
|
213
|
-
timeout=True,
|
|
214
|
-
idle_at_ts=ended_ts,
|
|
215
|
-
idle_seconds=idle_seconds,
|
|
216
|
-
terminated_by="idle-watchdog",
|
|
217
|
-
)
|
|
218
|
-
_write_status(status_path, status)
|
|
219
|
-
return exit_code
|
|
220
|
-
|
|
221
|
-
|
|
222
190
|
def main(argv: list[str]) -> int:
|
|
223
191
|
try:
|
|
224
|
-
invocation =
|
|
225
|
-
return
|
|
192
|
+
invocation = parse_invocation(argv[1:])
|
|
193
|
+
return run_worker(
|
|
194
|
+
invocation.strategy,
|
|
195
|
+
invocation.request,
|
|
196
|
+
presentation=invocation.presentation,
|
|
197
|
+
log_path=invocation.log_path,
|
|
198
|
+
status_path=invocation.status_path,
|
|
199
|
+
status_extra=invocation.status_extra,
|
|
200
|
+
)
|
|
226
201
|
except PreflightError as exc:
|
|
227
202
|
print(f"okstra-provider-exec: {exc}", file=sys.stderr)
|
|
228
203
|
return exc.exit_code
|
|
@@ -112,9 +112,26 @@ def _data_to_report_path(data_path: Path) -> Path:
|
|
|
112
112
|
return final_report_markdown_path(data_path)
|
|
113
113
|
|
|
114
114
|
|
|
115
|
+
def _project_id_of(task_key: str) -> str:
|
|
116
|
+
"""`project-id:task-group:task-id` 의 첫 세그먼트, 아니면 "".
|
|
117
|
+
|
|
118
|
+
A follow-up belongs to its parent's project by construction, so the new
|
|
119
|
+
task-key inherits that segment rather than re-deriving it from disk.
|
|
120
|
+
Omitting it is not a cosmetic difference: `okstra_project.state.
|
|
121
|
+
parse_task_key` raises on a two-segment key, and every catalog reader goes
|
|
122
|
+
through it, so one `<group>/<id>` entry took down `list_project_tasks` for
|
|
123
|
+
the whole project — the wizard could not render its first screen.
|
|
124
|
+
"""
|
|
125
|
+
parts = (task_key or "").split(":")
|
|
126
|
+
if len(parts) != 3 or not all(parts):
|
|
127
|
+
return ""
|
|
128
|
+
return parts[0]
|
|
129
|
+
|
|
130
|
+
|
|
115
131
|
def _spawn_one(
|
|
116
132
|
*,
|
|
117
133
|
project_root: Path,
|
|
134
|
+
project_id: str,
|
|
118
135
|
task_group: str,
|
|
119
136
|
parent_task_key: str,
|
|
120
137
|
parent_report_relative: str,
|
|
@@ -151,7 +168,7 @@ def _spawn_one(
|
|
|
151
168
|
origin = row["origin"].strip()
|
|
152
169
|
priority = (row.get("priority") or "P1").strip()
|
|
153
170
|
ticket_id = (row.get("ticketId") or "").strip()
|
|
154
|
-
new_task_key = f"{task_group}
|
|
171
|
+
new_task_key = f"{project_id}:{task_group}:{new_task_id}"
|
|
155
172
|
now = dt.datetime.now(dt.timezone.utc).isoformat()
|
|
156
173
|
|
|
157
174
|
spawned_meta: dict = {
|
|
@@ -257,6 +274,16 @@ def main(argv: list[str]) -> int:
|
|
|
257
274
|
print(f"data.json not found: {args.data_file}", file=sys.stderr)
|
|
258
275
|
return 1
|
|
259
276
|
|
|
277
|
+
project_id = _project_id_of(args.parent_task_key)
|
|
278
|
+
if not project_id:
|
|
279
|
+
print(
|
|
280
|
+
f"--parent-task-key must be project-id:task-group:task-id, got "
|
|
281
|
+
f"{args.parent_task_key!r} — a spawned task-key derived from it "
|
|
282
|
+
"would break every catalog read in this project.",
|
|
283
|
+
file=sys.stderr,
|
|
284
|
+
)
|
|
285
|
+
return 1
|
|
286
|
+
|
|
260
287
|
try:
|
|
261
288
|
data = json.loads(args.data_file.read_text(encoding="utf-8"))
|
|
262
289
|
except json.JSONDecodeError as exc:
|
|
@@ -297,6 +324,7 @@ def main(argv: list[str]) -> int:
|
|
|
297
324
|
continue
|
|
298
325
|
status, info = _spawn_one(
|
|
299
326
|
project_root=args.project_root,
|
|
327
|
+
project_id=project_id,
|
|
300
328
|
task_group=args.task_group,
|
|
301
329
|
parent_task_key=args.parent_task_key,
|
|
302
330
|
parent_report_relative=parent_report_relative,
|
|
@@ -2,13 +2,20 @@
|
|
|
2
2
|
#
|
|
3
3
|
# okstra-trace-cleanup.sh — close tmux panes created during okstra runs.
|
|
4
4
|
#
|
|
5
|
-
#
|
|
6
|
-
#
|
|
7
|
-
#
|
|
8
|
-
#
|
|
9
|
-
#
|
|
10
|
-
#
|
|
11
|
-
#
|
|
5
|
+
# Worker-compute panes are tmux-pane backend siblings. Their dispatcher tags
|
|
6
|
+
# each pane it owns with a pane-level user option (`@okstra_worker_run=<RUN_DIR>`),
|
|
7
|
+
# so panes are found server-wide by tag — no tmux env var or pane-id registry is
|
|
8
|
+
# needed, and the run-scoped tag keeps concurrent okstra runs from closing each
|
|
9
|
+
# other's panes.
|
|
10
|
+
#
|
|
11
|
+
# Trace panes were `tail -F` siblings the provider wrappers split, tagged
|
|
12
|
+
# `@okstra_trace_run` / `@okstra_status`. Those wrappers are now four-line
|
|
13
|
+
# entrypoints and worker progress renders into the worker's own pane, so
|
|
14
|
+
# NOTHING SPAWNS A TRACE PANE and neither tag has a writer left. The trace
|
|
15
|
+
# paths below still run and simply match nothing; `--reclaim-completed`, which
|
|
16
|
+
# keys on `@okstra_status`, is inert for the same reason. Kept rather than
|
|
17
|
+
# deleted because the hooks that call this script are already seeded on user
|
|
18
|
+
# machines — retiring the trace machinery is a deliberate follow-up.
|
|
12
19
|
#
|
|
13
20
|
# Two invocation shapes:
|
|
14
21
|
#
|
|
@@ -1,27 +1,34 @@
|
|
|
1
1
|
#!/usr/bin/env python3
|
|
2
|
-
"""okstra-wrapper-status.py —
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
|
|
6
|
-
`
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
2
|
+
"""okstra-wrapper-status.py — standalone CLI for the worker status sidecar.
|
|
3
|
+
|
|
4
|
+
A CLI worker dispatch runs a long-running CLI in the background, and the
|
|
5
|
+
polling channel that watches it carries only stdout plus a binary
|
|
6
|
+
`running`/`completed` state. Several recovery decisions need more —
|
|
7
|
+
specifically, "did this worker start at all, when, and how did it finish?" — so
|
|
8
|
+
a small JSON sidecar is written at `<prompt-path>.status.json` that survives
|
|
9
|
+
independent of the polling channel.
|
|
10
|
+
|
|
11
|
+
**Nothing calls this script.** Every provider entrypoint now runs through
|
|
12
|
+
`scripts/okstra_ctl/worker_runner.py`, which writes the same document
|
|
13
|
+
in-process; the shell wrappers that shelled out to this file were its only
|
|
14
|
+
callers and they are gone. It is kept for now rather than deleted because
|
|
15
|
+
removing it also means touching the build sync list, the install payload, its
|
|
16
|
+
own test module and the docs that name it — a deliberate follow-up, not a side
|
|
17
|
+
effect of the wrapper migration. `scripts/okstra_ctl/wrapper_status.py` is the
|
|
18
|
+
reader, and what it reads is what the runner writes.
|
|
19
|
+
|
|
20
|
+
Consumers of the sidecar:
|
|
21
|
+
|
|
22
|
+
* A CLI worker's diagnostic step: read `log_path` to capture a tail when
|
|
23
|
+
`exit_code == 0` but the canonical Result file is absent.
|
|
24
|
+
* Lead: cross-check `started_ts` / `ended_ts` to distinguish "worker hung
|
|
18
25
|
before CLI launched" from "CLI finished but never wrote artifact" when
|
|
19
26
|
applying the redispatch policy (see team-contract "Lead Redispatch
|
|
20
27
|
Policy on Result-Missing").
|
|
21
28
|
|
|
22
|
-
Failures are deliberately non-fatal for the caller — the
|
|
23
|
-
|
|
24
|
-
|
|
29
|
+
Failures are deliberately non-fatal for the caller — the caller's main job is
|
|
30
|
+
to run the underlying CLI; a missing sidecar must not break that. On any error
|
|
31
|
+
the script prints a one-line diagnostic to stderr and exits 0.
|
|
25
32
|
|
|
26
33
|
Schema (schemaVersion 1):
|
|
27
34
|
|
|
@@ -57,7 +57,7 @@ Use the screen to tell "still working" from "stuck", and to see at a glance whic
|
|
|
57
57
|
- Worker completion is valid only from `workerDispatches[]`, terminal status sidecars, and required Result Paths. Pane creation alone is not completion.
|
|
58
58
|
- Reverify uses a fresh jobs file at `runs/<task-type>/state/reverify-jobs-r<N>-<task-type>-<seq>.json`, sets `dispatchKind: "reverify-r<N>"`, and dispatches with `okstra team dispatch --project-root <root> --run-manifest <path> --dispatch-kind reverify-r<N> --jobs-file <jobs-file>`.
|
|
59
59
|
- Report-writer uses a fresh one-job jobs file with `dispatchKind: "report-writer"` and the same schema, then dispatches through `okstra team dispatch --project-root <root> --run-manifest <path> --jobs-file <jobs-file>`.
|
|
60
|
-
- Every reverify or report-writer jobs file carries `workerId`, `provider`, `role`, `modelExecutionValue`, `promptPath`, `resultPath`, `workerResultPath`, and `completionPaths`. For reverify, set `role` to `worker-reverify-r<N>`
|
|
60
|
+
- Every reverify or report-writer jobs file carries `workerId`, `provider`, `role`, `modelExecutionValue`, `promptPath`, `resultPath`, `workerResultPath`, and `completionPaths`. For reverify, set `role` to `worker-reverify-r<N>` — the role selects the dispatch's idle budget and is recorded in the run's status sidecar, so it must name the actual assignment. The report-writer completion paths include both data.json and the worker-results audit file.
|
|
61
61
|
- After either dispatch, run `okstra team await --project-root <root> --run-manifest <path>` before evaluating terminal status or completion paths.
|
|
62
62
|
|
|
63
63
|
## Completion, cleanup, and resume
|
|
@@ -540,7 +540,7 @@ Schema rules:
|
|
|
540
540
|
|
|
541
541
|
## Coverage critic pass
|
|
542
542
|
|
|
543
|
-
Runs only when `convergence.critic.enabled == true` (set by `--critic <provider>` or the okstra-run `critic_pick` step; default off). Applies to the three finding-producing phases (`requirements-discovery`, `error-analysis`, `implementation-planning`); for `final-verification` the critic runs in a different mode — see §"Acceptance critic pass (final-verification)". This pass targets **
|
|
543
|
+
Runs only when `convergence.critic.enabled == true` (set by `--critic <provider>` or the okstra-run `critic_pick` step; default off). Applies to the three finding-producing phases (`requirements-discovery`, `error-analysis`, `implementation-planning`); for `final-verification` the critic runs in a different mode — see §"Acceptance critic pass (final-verification)". This pass targets **scope in both directions** — findings that are missing (coverage) and work the findings propose that no requirement asked for (over-scope) — distinct from convergence, which targets **agreement quality** among the findings already raised. The pass keeps its `coverage` mode id and `gaps` vocabulary for both halves; the two are told apart by each candidate's `category`, so no schema or reducer distinguishes them.
|
|
544
544
|
|
|
545
545
|
### When
|
|
546
546
|
|
|
@@ -554,10 +554,12 @@ Dispatch one fresh pass to `config.critic.provider` through `redispatch_worker`,
|
|
|
554
554
|
|
|
555
555
|
The `-worker-` token is load-bearing, not decoration: the critic prompt carries the same generated anchor headers as every other worker ([team-contract](./team-contract.md) §"Worker prompts"), and its `**Audit sidecar path:**` comes from passing that result path through `okstra_ctl.worker_artifact_paths.audit_sidecar_rel()`, which inserts `-audit-` after the token and raises without it. A `<provider>-critic-...` name leaves the lead choosing between breaking the contract and hand-inventing the sidecar name. Note that `originWorker` stays `"<provider>-critic"` — that is a worker id in the convergence state, not a filename, and the two do not have to match.
|
|
556
556
|
|
|
557
|
-
The critic prompt
|
|
557
|
+
The critic prompt carries the full analysis contract those anchor headers belong to — worker preamble, error-contract path, audit sidecar, packet boundary — but it is **exempt from the initial-analysis equality group**. Initial analysis workers must receive byte-identical normalized bodies (`worker_prompt_contract.validate_analysis_prompt_set`), and a critic body is deliberately unlike them, so `okstra_ctl.worker_prompt_policy.resolve_prompt_plan` resolves `dispatchKind = "critic"` to `equality_group = None` while keeping `audience = "analysis"`. Never reshape a critic prompt to match the initial one to satisfy that check — matching it would delete the pass. **Enforced:** `tests/contract/test_critic_prompt_equality_exemption.py` pins both directions — the critic is exempt, and two mismatched *initial* prompts still fail.
|
|
558
558
|
|
|
559
|
-
|
|
560
|
-
|
|
559
|
+
The critic prompt seeds the consolidated findings and asks for two things only — coverage gaps and unrequested work. The second half exists because every other scope check in the lifecycle runs *later*: the plan-body `P-Opt-*` verdict judges options this phase has not authored yet (report-writer writes the plan in Phase 6), and the `implementation` verifier judges a diff. Here the unit is a **finding**, which is the input those stages build on — catching unrequested work while it is still a finding is the cheapest place to catch it at all.
|
|
560
|
+
|
|
561
|
+
Required reading before proposing a gap or an over-scope candidate:
|
|
562
|
+
- the current run's `analysis-packet.md` for requirements and phase scope — for the over-scope half this is the authority you search against, so read it before judging any finding unrequested;
|
|
561
563
|
- `convergence-groups-<task-type>-<seq>.json` for the complete Round 0 ledger;
|
|
562
564
|
- every initial analysis-worker result named by team-state;
|
|
563
565
|
- each matching audit sidecar, to distinguish an uninspected path from a claim that was inspected but summarized during grouping.
|
|
@@ -565,13 +567,30 @@ Required reading before proposing a gap:
|
|
|
565
567
|
Operational guardrails are not task requirements. A gap must trace to a brief requirement, an analysis-packet scope item, a source path the packet authorizes, or an evidence claim in a worker result. Do NOT infer missing verification from a one-line summary; open the named result and audit sidecar first.
|
|
566
568
|
|
|
567
569
|
```
|
|
568
|
-
You are the
|
|
569
|
-
|
|
570
|
+
You are the scope critic for <task-key>. Below are the consolidated findings the
|
|
571
|
+
workers produced. Your job has exactly two halves. Answer both.
|
|
572
|
+
|
|
573
|
+
(1) MISSING — name what nobody covered:
|
|
570
574
|
- files / directories / execution paths nobody inspected,
|
|
571
575
|
- requirements or acceptance points with zero findings,
|
|
572
576
|
- claims raised but never verified.
|
|
573
|
-
For each
|
|
574
|
-
|
|
577
|
+
For each, emit a NEW finding with evidence (file:line or the requirement quote).
|
|
578
|
+
|
|
579
|
+
(2) UNREQUESTED — name work these findings propose that no requirement asked for:
|
|
580
|
+
- a finding whose proposed change serves no requirement, scope item, or
|
|
581
|
+
acceptance point you can QUOTE from the analysis packet,
|
|
582
|
+
- an abstraction, configuration knob, or generalization proposed for a caller or
|
|
583
|
+
a case nobody has stated,
|
|
584
|
+
- a rewrite, migration, or cleanup of code the requirements never mention.
|
|
585
|
+
For each, emit a candidate with `category: "unrequested-scope"`, quote the
|
|
586
|
+
proposed work verbatim, and state which requirement you searched for and did not
|
|
587
|
+
find.
|
|
588
|
+
|
|
589
|
+
Do NOT restate an existing finding. Judge (2) against the analysis packet's
|
|
590
|
+
requirements and scope, never against your own preference for how the code should
|
|
591
|
+
look — "I would have done it differently" is not unrequested work, and neither is
|
|
592
|
+
work the packet authorizes but you consider unnecessary. If a half has nothing,
|
|
593
|
+
say so explicitly for that half; silence on one half is an incomplete result.
|
|
575
594
|
```
|
|
576
595
|
|
|
577
596
|
### Gap verification (1 adversarial reverify round)
|
|
@@ -579,8 +598,17 @@ Each critic gap enters the verification queue as a finding with `originWorker =
|
|
|
579
598
|
|
|
580
599
|
**A gap that received no verdict is NOT a rejected gap (BLOCKING).** Dropping applies only to gaps the voters actually judged. A gap can also end the round *unjudged* — the verification dispatch returned a terminal non-result (`timeout`, `error`, no result file), the returned result covered only some of the gaps, or no non-critic analyser was available to vote at all. Nobody inspected those, so classifying them as hallucinations is a fabricated verdict. Each one MUST be recorded as a `## 5. Missing Information and Risks` row (`missingInformation`, `source: "critic-unverified"`) whose `risk` names the gap and the reason verification did not complete, and counted in `config.critic.gapsUnverified`. They are **not** promoted to findings (unverified) and **not** raised as `clarification` items — an unverified gap needs an analyser to verify it on the next run, not a decision from the user. Silently losing them is a contract violation: the batch that times out is exactly the batch of gaps too expensive to check, so the highest-risk items are the ones that vanish.
|
|
581
600
|
|
|
601
|
+
**`category: "unrequested-scope"` candidates are classified the same way but disposed of differently.** A coverage gap the voters contest is a hallucination — nothing was actually missing, so dropping it costs one wasted verification. An over-scope candidate the voters contest is a *disagreement about whether the work was asked for*, and dropping that silently returns the run to the state this half exists to change. So:
|
|
602
|
+
|
|
603
|
+
- `full-consensus` / `partial-consensus` → merge as a finding, exactly like a coverage gap. The merged finding names the unrequested work and the requirement search that came up empty; the phase's own deliverable rules decide whether it lands as a dropped item or a clarification row.
|
|
604
|
+
- `contested` / `worker-unique` → counted in `config.critic.gapsRejected` (unchanged accounting) but ALSO recorded as a `## 5. Missing Information and Risks` row with `source: "critic-unconfirmed-scope"`, whose `risk` quotes the proposed work and the split verdict. It is **not** promoted to a finding: a contested over-scope claim must not block a plan on one worker's taste. The user reads the row and decides.
|
|
605
|
+
- no verdict at all → the `critic-unverified` rule above applies unchanged.
|
|
606
|
+
|
|
607
|
+
The asymmetry is deliberate and runs the opposite way from the coverage half: a false "you missed something" costs a verification, while a false "you built too much" costs a real requirement — so the first may be dropped outright and the second is recorded either way. This is the same reasoning that makes the `final-verification` acceptance critic never drop a candidate, applied to the one direction where dropping is otherwise the default.
|
|
608
|
+
|
|
582
609
|
### State
|
|
583
610
|
- `convergence.critic` manifest block: `{ enabled, provider, modelExecutionValue }`.
|
|
611
|
+
- Each candidate's `category` tells the two halves apart: literal `"unrequested-scope"` for the over-scope half, any other value for a coverage gap. `schemas/convergence-critic-results-v1.0.schema.json` leaves `category` a free string, so this needs no schema or reducer change — but it also means nothing machine-checks the spelling. A misspelled category is read as a coverage gap and silently takes the drop-on-contested path.
|
|
584
612
|
- The lead passes one canonical coverage batch with `{ schemaVersion, taskKey, mode, provider, modelExecutionValue, dispatches, gaps }`; each gap carries its candidate fields plus `gapId` and `votes`. `dispatches[]` contains exactly one row for every Phase 4 analyser except the critic, even when execution did not produce a result: persist `status: timeout | error | not-run` and the elapsed `durationMs` instead of omitting that analyser. `apply-critic-gaps` rejects a non-terminal main queue, duplicate or missing analysers, unknown workers, critic dispatches/votes, votes without a completed dispatch, and a second batch.
|
|
585
613
|
- Convergence state artifact: merged gaps appear in `findings[]` with `source: "critic"` and `rounds: []`. The separate `criticVerification.gaps[]` ledger retains each gap's `summary`, `category`, `ticketIds`, `originEvidence`, optional `evidenceArtifacts`, classification, merge link, and votes. Strict v1.3 validation deterministically replays each complete critic-origin finding from that ledger; a critic batch never increments `roundHistory` or `totalRounds` and never creates a fake main round.
|
|
586
614
|
- `config.critic` is `{ provider, modelExecutionValue, gapsProposed, gapsMerged, gapsRejected, gapsUnverified }`, with `gapsProposed = gapsMerged + gapsRejected + gapsUnverified`. `full-consensus` / `partial-consensus` gaps merge, `contested` / `worker-unique` gaps count as rejected, and gaps with no usable analyser vote appear in both the ledger and final `unverifiedGaps[]`.
|