okstra 0.164.0 → 0.165.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (84) hide show
  1. package/README.md +1 -1
  2. package/docs/architecture.md +12 -8
  3. package/docs/cli.md +7 -3
  4. package/docs/for-ai/README.md +2 -2
  5. package/docs/for-ai/skills/okstra-inspect.md +2 -2
  6. package/docs/for-ai/skills/okstra-user-response.md +2 -2
  7. package/docs/project-structure-overview.md +15 -9
  8. package/package.json +1 -1
  9. package/runtime/BUILD.json +2 -2
  10. package/runtime/agents/workers/antigravity-worker.md +9 -7
  11. package/runtime/agents/workers/codex-worker.md +9 -7
  12. package/runtime/agents/workers/grok-worker.md +6 -4
  13. package/runtime/agents/workers/kimi-worker.md +6 -4
  14. package/runtime/bin/okstra-antigravity-exec.sh +1 -340
  15. package/runtime/bin/okstra-claude-exec.sh +1 -178
  16. package/runtime/bin/okstra-codex-exec.sh +1 -467
  17. package/runtime/bin/okstra-provider-exec.py +165 -190
  18. package/runtime/bin/okstra-trace-cleanup.sh +14 -7
  19. package/runtime/bin/okstra-wrapper-status.py +26 -19
  20. package/runtime/prompts/lead/adapters/cmux.md +1 -1
  21. package/runtime/prompts/lead/convergence.md +36 -8
  22. package/runtime/prompts/lead/okstra-lead-contract.md +23 -1
  23. package/runtime/prompts/lead/plan-body-verification.md +9 -1
  24. package/runtime/prompts/lead/report-writer.md +1 -0
  25. package/runtime/prompts/lead/team-contract.md +3 -3
  26. package/runtime/prompts/profiles/_common-contract.md +9 -1
  27. package/runtime/prompts/profiles/_coverage-critic.md +1 -1
  28. package/runtime/prompts/profiles/_implementation-diff-review.md +3 -1
  29. package/runtime/prompts/profiles/_implementation-self-check.md +1 -1
  30. package/runtime/prompts/profiles/_implementation-verifier.md +3 -1
  31. package/runtime/prompts/profiles/implementation-planning.md +5 -3
  32. package/runtime/python/okstra_ctl/adapters/hosts/claude-code/relay.md +1 -1
  33. package/runtime/python/okstra_ctl/adapters/hosts/external/relay.md +1 -1
  34. package/runtime/python/okstra_ctl/adapters/providers/antigravity/adapter.py +148 -0
  35. package/runtime/python/okstra_ctl/adapters/providers/claude/adapter.py +55 -0
  36. package/runtime/python/okstra_ctl/adapters/providers/codex/adapter.py +41 -0
  37. package/runtime/python/okstra_ctl/adapters/providers/grok/adapter.py +44 -0
  38. package/runtime/python/okstra_ctl/adapters/providers/kimi/adapter.py +42 -0
  39. package/runtime/python/okstra_ctl/dispatch_core.py +5 -1
  40. package/runtime/python/okstra_ctl/dispatch_state.py +10 -0
  41. package/runtime/python/okstra_ctl/domain/provider.py +5 -1
  42. package/runtime/python/okstra_ctl/domain/worker_exec.py +102 -0
  43. package/runtime/python/okstra_ctl/domain/worker_role.py +34 -0
  44. package/runtime/python/okstra_ctl/domain/worker_stream.py +261 -0
  45. package/runtime/python/okstra_ctl/incremental_scope.py +16 -4
  46. package/runtime/python/okstra_ctl/report_html/common.py +71 -25
  47. package/runtime/python/okstra_ctl/report_html/models.py +5 -0
  48. package/runtime/python/okstra_ctl/report_html/render.py +1 -1
  49. package/runtime/python/okstra_ctl/report_html/run_usage.py +19 -0
  50. package/runtime/python/okstra_ctl/report_html/view_models/implementation_planning.py +14 -0
  51. package/runtime/python/okstra_ctl/report_views.py +44 -16
  52. package/runtime/python/okstra_ctl/stage_citations.py +52 -15
  53. package/runtime/python/okstra_ctl/user_response.py +45 -29
  54. package/runtime/python/okstra_ctl/wizard.py +13 -9
  55. package/runtime/python/okstra_ctl/worker_prompt_policy.py +10 -3
  56. package/runtime/python/okstra_ctl/worker_request.py +140 -0
  57. package/runtime/python/okstra_ctl/worker_runner.py +622 -0
  58. package/runtime/python/okstra_token_usage/collect.py +8 -1
  59. package/runtime/python/okstra_token_usage/report.py +42 -0
  60. package/runtime/python/okstra_token_usage/task_totals.py +88 -0
  61. package/runtime/schemas/final-report-v1.0.schema.json +70 -0
  62. package/runtime/schemas/final-report-v2.0.schema.json +90 -0
  63. package/runtime/skills/okstra-inspect/SKILL.md +1 -2
  64. package/runtime/skills/okstra-inspect/facets/logs.md +5 -5
  65. package/runtime/skills/okstra-inspect/facets/run-audit.md +3 -3
  66. package/runtime/skills/okstra-run/SKILL.md +1 -1
  67. package/runtime/skills/okstra-user-response/SKILL.md +15 -5
  68. package/runtime/templates/report-writer-prompt-preamble.md +1 -0
  69. package/runtime/templates/reports/html/assets/base.css +8 -4
  70. package/runtime/templates/reports/html/base.template.html +12 -6
  71. package/runtime/templates/reports/html/i18n/en.json +29 -6
  72. package/runtime/templates/reports/html/i18n/ko.json +29 -6
  73. package/runtime/templates/reports/html/macros/forms.html +9 -3
  74. package/runtime/templates/reports/html/tasks/implementation-planning.template.html +14 -19
  75. package/runtime/templates/reports/report.js +59 -26
  76. package/runtime/templates/reports/user-response.template.md +12 -8
  77. package/runtime/validators/validate-run.py +88 -7
  78. package/runtime/validators/validate_session_conformance.py +62 -1
  79. package/src/cli-registry.mjs +0 -7
  80. package/runtime/bin/okstra-wrapper-agy-stream.py +0 -61
  81. package/runtime/python/okstra_ctl/error_issue.py +0 -640
  82. package/runtime/python/okstra_ctl/issue_signals.py +0 -186
  83. package/runtime/skills/okstra-inspect/facets/error-issue.md +0 -77
  84. package/src/commands/inspect/error-issue.mjs +0 -27
@@ -1,121 +1,184 @@
1
1
  #!/usr/bin/env python3
2
- """Run an external LLM CLI with the shared okstra wrapper contract."""
2
+ """Run an external LLM CLI with the shared okstra wrapper contract.
3
+
4
+ Argument parsing and provider lookup only — the run itself belongs to
5
+ ``okstra_ctl.worker_runner``, which every provider shares. What stays here is
6
+ the request the provider strategies then take on trust: resolved paths, the
7
+ write scope in the order the CLIs are told it, and the role's idle budget.
8
+ """
3
9
  from __future__ import annotations
4
10
 
5
- import json
6
11
  import os
7
- import selectors
8
12
  import shutil
9
- import signal
10
- import subprocess
11
13
  import sys
12
- import time
13
14
  from dataclasses import dataclass
14
15
  from pathlib import Path
15
- from typing import Callable
16
+ from typing import Any, Mapping
17
+
18
+ _HERE = Path(__file__).resolve().parent
19
+ # ``okstra_ctl`` sits beside this file in the repo (``scripts/``) but under
20
+ # ``~/.okstra/lib/python/`` once installed, and the four-line shell entrypoints
21
+ # that exec this script set no PYTHONPATH. Offer both, repo first.
22
+ _HOME_LIB = (
23
+ Path(os.environ.get("OKSTRA_HOME", str(Path.home() / ".okstra"))) / "lib" / "python"
24
+ )
25
+ sys.path.insert(0, str(_HERE))
26
+ if _HOME_LIB.is_dir() and str(_HOME_LIB) not in sys.path:
27
+ sys.path.append(str(_HOME_LIB))
28
+
29
+ from okstra_ctl.domain.provider import ProviderSpec, UnknownProviderError # noqa: E402
30
+ from okstra_ctl.domain.worker_exec import ( # noqa: E402
31
+ ExecutionStrategy,
32
+ WorkerExecRequest,
33
+ )
34
+ from okstra_ctl.registry.provider_registry import ( # noqa: E402
35
+ default_provider_registry,
36
+ )
37
+ from okstra_ctl.worker_request import build_request, idle_timeout # noqa: E402
38
+ from okstra_ctl.worker_runner import LIVE, QUIET, run_worker # noqa: E402
39
+
40
+ _USAGE = (
41
+ "usage: okstra-provider-exec.py <provider> <project-root> "
42
+ "<model-execution-value> <prompt-path> [worktree-path] [role] "
43
+ "[idle-timeout-seconds] [--presentation live|quiet]"
44
+ )
45
+
46
+ _PRESENTATION_FLAG = "--presentation"
47
+ _PRESENTATIONS = (LIVE, QUIET)
48
+
49
+
50
+ class PreflightError(Exception):
51
+ def __init__(self, exit_code: int, message: str) -> None:
52
+ super().__init__(message)
53
+ self.exit_code = exit_code
16
54
 
17
55
 
18
56
  @dataclass(frozen=True)
19
- class ProviderCommand:
20
- binary: str
21
- wrapper: str
22
- build_args: Callable[[str, str, str], list[str]]
57
+ class Invocation:
58
+ strategy: ExecutionStrategy
59
+ request: WorkerExecRequest
60
+ presentation: str
61
+ log_path: Path
62
+ status_path: Path
63
+ status_extra: Mapping[str, Any]
64
+
65
+
66
+ def parse_invocation(argv: list[str]) -> Invocation:
67
+ """Resolve the wrapper's positional contract into one runnable dispatch."""
68
+ positional, presentation = _take_presentation(argv)
69
+ if not 4 <= len(positional) <= 7:
70
+ raise PreflightError(64, _USAGE)
71
+ provider_id, project_root_raw, model, prompt_raw = positional[:4]
72
+ worktree_raw = positional[4] if len(positional) >= 5 else ""
73
+ role = positional[5] if len(positional) >= 6 and positional[5] else "worker"
74
+ timeout_raw = positional[6] if len(positional) >= 7 and positional[6] else ""
75
+
76
+ spec = _provider_spec(provider_id)
77
+ project_root = _existing_dir(project_root_raw, 65, "project-root")
78
+ if not model:
79
+ raise PreflightError(66, "model-execution-value is empty")
80
+ prompt_path = _existing_file(prompt_raw, 67, "prompt-path")
81
+ idle_timeout_seconds = _idle_timeout(timeout_raw, role)
82
+ worktree = (
83
+ _existing_dir(worktree_raw, 68, "worktree-path") if worktree_raw else None
84
+ )
23
85
 
86
+ request = build_request(
87
+ prompt_text=prompt_path.read_text(encoding="utf-8"),
88
+ model=model,
89
+ project_root=project_root,
90
+ worktree_path=worktree,
91
+ role=role,
92
+ idle_timeout_seconds=idle_timeout_seconds,
93
+ )
94
+ strategy = spec.exec_strategy
95
+ _check_command(strategy, request)
96
+ return Invocation(
97
+ strategy=strategy,
98
+ request=request,
99
+ presentation=presentation,
100
+ log_path=_log_path(prompt_path),
101
+ status_path=Path(f"{prompt_path}.status.json"),
102
+ status_extra={"wrapper": spec.wrapper, "role": role},
103
+ )
24
104
 
25
- def _grok_args(prompt: str, model: str, cwd: str) -> list[str]:
26
- return [
27
- "grok",
28
- "-p",
29
- prompt,
30
- "-m",
31
- model,
32
- "--output-format",
33
- "streaming-json",
34
- "--cwd",
35
- cwd,
36
- ]
37
105
 
106
+ def _take_presentation(argv: list[str]) -> tuple[list[str], str]:
107
+ """Split the one flag out of an otherwise positional argv.
108
+
109
+ Defaults to ``quiet``. ``live`` is only ever right where a screen was
110
+ declared, and the only callers that can declare one are the pane backends —
111
+ which pass the flag explicitly. Defaulting the other way assumed a screen
112
+ that a subagent dispatch does not have, and sent every worker's progress
113
+ into its caller's context window instead.
114
+ """
115
+ positional: list[str] = []
116
+ presentation = QUIET
117
+ index = 0
118
+ while index < len(argv):
119
+ if argv[index] != _PRESENTATION_FLAG:
120
+ positional.append(argv[index])
121
+ index += 1
122
+ continue
123
+ if index + 1 >= len(argv):
124
+ raise PreflightError(64, f"{_PRESENTATION_FLAG} needs a value: {_USAGE}")
125
+ presentation = argv[index + 1]
126
+ index += 2
127
+ if presentation not in _PRESENTATIONS:
128
+ allowed = " | ".join(_PRESENTATIONS)
129
+ raise PreflightError(
130
+ 64, f"unsupported presentation {presentation!r}. Allowed values: {allowed}"
131
+ )
132
+ return positional, presentation
38
133
 
39
- def _kimi_args(prompt: str, model: str, _cwd: str) -> list[str]:
40
- return ["kimi", "-p", prompt, "-m", model, "--output-format", "stream-json"]
41
134
 
135
+ def _provider_spec(provider_id: str) -> ProviderSpec:
136
+ try:
137
+ spec = default_provider_registry().resolve(provider_id)
138
+ except UnknownProviderError as exc:
139
+ raise PreflightError(64, str(exc)) from exc
140
+ if spec.exec_strategy is None:
141
+ raise PreflightError(
142
+ 64, f"provider {provider_id!r} has no execution strategy to run"
143
+ )
144
+ return spec
42
145
 
43
- PROVIDERS = {
44
- "grok": ProviderCommand("grok", "okstra-grok-exec.sh", _grok_args),
45
- "kimi": ProviderCommand("kimi", "okstra-kimi-exec.sh", _kimi_args),
46
- }
47
146
 
147
+ def _existing_dir(raw: str, exit_code: int, label: str) -> Path:
148
+ """Existence only — `build_request` owns the resolving."""
149
+ path = Path(raw) if raw else None
150
+ if path is None or not path.is_dir():
151
+ raise PreflightError(
152
+ exit_code, f"{label} is missing or not a directory: {raw!r}"
153
+ )
154
+ return path
48
155
 
49
- @dataclass(frozen=True)
50
- class Invocation:
51
- provider: ProviderCommand
52
- project_root: Path
53
- model: str
54
- prompt_path: Path
55
- execution_root: Path
56
- role: str
57
- idle_timeout_seconds: int
58
156
 
157
+ def _existing_file(raw: str, exit_code: int, label: str) -> Path:
158
+ path = Path(raw) if raw else None
159
+ if path is None or not path.is_file():
160
+ raise PreflightError(exit_code, f"{label} is missing or not a file: {raw!r}")
161
+ return path.resolve()
59
162
 
60
- class PreflightError(Exception):
61
- def __init__(self, exit_code: int, message: str) -> None:
62
- super().__init__(message)
63
- self.exit_code = exit_code
64
163
 
164
+ def _idle_timeout(raw: str, role: str) -> int:
165
+ try:
166
+ return idle_timeout(raw, role)
167
+ except ValueError as exc:
168
+ raise PreflightError(69, str(exc)) from exc
65
169
 
66
- def _parse_invocation(argv: list[str]) -> Invocation:
67
- if len(argv) < 4 or len(argv) > 7:
68
- raise PreflightError(
69
- 64,
70
- "usage: okstra-provider-exec.py <provider> <project-root> <model-execution-value> "
71
- "<prompt-path> [worktree-path] [role] [idle-timeout-seconds]",
72
- )
73
- provider_id, project_root_raw, model, prompt_raw = argv[:4]
74
- provider = PROVIDERS.get(provider_id)
75
- if provider is None:
76
- raise PreflightError(64, f"unsupported provider: {provider_id}")
77
- worktree_raw = argv[4] if len(argv) >= 5 else ""
78
- role = argv[5] if len(argv) >= 6 and argv[5] else "worker"
79
- default_timeout = 1500 if role in {"executor", "verifier"} else 600
80
- timeout_raw = argv[6] if len(argv) >= 7 else str(default_timeout)
81
- return _validate_invocation(
82
- provider, project_root_raw, model, prompt_raw, worktree_raw, role, timeout_raw
83
- )
84
170
 
171
+ def _check_command(strategy: ExecutionStrategy, request: WorkerExecRequest) -> None:
172
+ """Refuse a missing CLI before the run leaves any artifact behind.
85
173
 
86
- def _validate_invocation(
87
- provider: ProviderCommand,
88
- project_root_raw: str,
89
- model: str,
90
- prompt_raw: str,
91
- worktree_raw: str,
92
- role: str,
93
- timeout_raw: str,
94
- ) -> Invocation:
95
- project_root = Path(project_root_raw)
96
- prompt_path = Path(prompt_raw)
97
- if not project_root_raw or not project_root.is_dir():
98
- raise PreflightError(65, f"project-root is missing or not a directory: {project_root_raw!r}")
99
- if not model:
100
- raise PreflightError(66, "model-execution-value is empty")
101
- if not prompt_raw or not prompt_path.is_file():
102
- raise PreflightError(67, f"prompt-path is missing or not a file: {prompt_raw!r}")
103
- if not timeout_raw.isdigit():
104
- raise PreflightError(69, f"idle-timeout-seconds must be a non-negative integer: {timeout_raw!r}")
105
- execution_root = Path(worktree_raw) if worktree_raw else project_root
106
- if worktree_raw and not execution_root.is_dir():
107
- raise PreflightError(68, f"worktree-path was provided but is not a directory: {worktree_raw!r}")
108
- if shutil.which(provider.binary) is None:
109
- raise PreflightError(127, f"{provider.binary} CLI is not installed on PATH")
110
- return Invocation(
111
- provider=provider,
112
- project_root=project_root.resolve(),
113
- model=model,
114
- prompt_path=prompt_path.resolve(),
115
- execution_root=execution_root.resolve(),
116
- role=role,
117
- idle_timeout_seconds=int(timeout_raw),
118
- )
174
+ Building the command is the only truthful way to learn which binary this
175
+ provider runs, and it is a pure call. The check stays out of the runner so a
176
+ refused dispatch leaves no `started` status sidecar for the liveness probe to
177
+ read as a worker that launched.
178
+ """
179
+ binary = strategy.build_command(request).argv[0]
180
+ if shutil.which(binary) is None:
181
+ raise PreflightError(127, f"{binary} CLI is not installed on PATH")
119
182
 
120
183
 
121
184
  def _log_path(prompt_path: Path) -> Path:
@@ -124,105 +187,17 @@ def _log_path(prompt_path: Path) -> Path:
124
187
  return Path(f"{prompt_path}.log")
125
188
 
126
189
 
127
- def _write_status(path: Path, status: dict[str, object]) -> None:
128
- temporary = Path(f"{path}.tmp")
129
- temporary.write_text(json.dumps(status, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
130
- os.replace(temporary, path)
131
-
132
-
133
- def _terminate_process(process: subprocess.Popen[bytes]) -> None:
134
- try:
135
- os.killpg(process.pid, signal.SIGTERM)
136
- except ProcessLookupError:
137
- return
138
- try:
139
- process.wait(timeout=5)
140
- except subprocess.TimeoutExpired:
141
- try:
142
- os.killpg(process.pid, signal.SIGKILL)
143
- except ProcessLookupError:
144
- pass
145
- process.wait()
146
-
147
-
148
- def _stream_process(
149
- process: subprocess.Popen[bytes], log_file, idle_timeout_seconds: int
150
- ) -> tuple[int, bool, int]:
151
- selector = selectors.DefaultSelector()
152
- assert process.stdout is not None
153
- selector.register(process.stdout, selectors.EVENT_READ)
154
- last_output = time.monotonic()
155
- timed_out = False
156
- idle_seconds = 0
157
- while selector.get_map():
158
- for key, _ in selector.select(timeout=0.25):
159
- chunk = os.read(key.fd, 8192)
160
- if not chunk:
161
- selector.unregister(key.fileobj)
162
- continue
163
- last_output = time.monotonic()
164
- sys.stdout.buffer.write(chunk)
165
- sys.stdout.buffer.flush()
166
- log_file.write(chunk)
167
- log_file.flush()
168
- idle_seconds = int(time.monotonic() - last_output)
169
- if idle_timeout_seconds and idle_seconds >= idle_timeout_seconds and process.poll() is None:
170
- timed_out = True
171
- _terminate_process(process)
172
- exit_code = process.wait()
173
- return (124 if timed_out else exit_code), timed_out, idle_seconds
174
-
175
-
176
- def _run(invocation: Invocation) -> int:
177
- prompt = invocation.prompt_path.read_text(encoding="utf-8")
178
- command = invocation.provider.build_args(prompt, invocation.model, str(invocation.execution_root))
179
- status_path = Path(f"{invocation.prompt_path}.status.json")
180
- log_path = _log_path(invocation.prompt_path)
181
- started_ts = int(time.time())
182
- started_monotonic = time.monotonic()
183
- status: dict[str, object] = {
184
- "schemaVersion": 1,
185
- "wrapper": invocation.provider.wrapper,
186
- "role": invocation.role,
187
- "pid": os.getpid(),
188
- "started_ts": started_ts,
189
- "log_path": str(log_path),
190
- "stage": "started",
191
- }
192
- _write_status(status_path, status)
193
- with log_path.open("wb") as log_file:
194
- process = subprocess.Popen(
195
- command,
196
- cwd=invocation.execution_root,
197
- stdout=subprocess.PIPE,
198
- stderr=subprocess.STDOUT,
199
- start_new_session=True,
200
- )
201
- exit_code, timed_out, idle_seconds = _stream_process(
202
- process, log_file, invocation.idle_timeout_seconds
203
- )
204
- ended_ts = int(time.time())
205
- status.update(
206
- stage="exited",
207
- exit_code=exit_code,
208
- ended_ts=ended_ts,
209
- duration_ms=int((time.monotonic() - started_monotonic) * 1000),
210
- )
211
- if timed_out:
212
- status.update(
213
- timeout=True,
214
- idle_at_ts=ended_ts,
215
- idle_seconds=idle_seconds,
216
- terminated_by="idle-watchdog",
217
- )
218
- _write_status(status_path, status)
219
- return exit_code
220
-
221
-
222
190
  def main(argv: list[str]) -> int:
223
191
  try:
224
- invocation = _parse_invocation(argv[1:])
225
- return _run(invocation)
192
+ invocation = parse_invocation(argv[1:])
193
+ return run_worker(
194
+ invocation.strategy,
195
+ invocation.request,
196
+ presentation=invocation.presentation,
197
+ log_path=invocation.log_path,
198
+ status_path=invocation.status_path,
199
+ status_extra=invocation.status_extra,
200
+ )
226
201
  except PreflightError as exc:
227
202
  print(f"okstra-provider-exec: {exc}", file=sys.stderr)
228
203
  return exc.exit_code
@@ -2,13 +2,20 @@
2
2
  #
3
3
  # okstra-trace-cleanup.sh — close tmux panes created during okstra runs.
4
4
  #
5
- # Trace panes are `tail -F` siblings spawned by the codex/antigravity wrappers
6
- # (`okstra-codex-exec.sh`, `okstra-antigravity-exec.sh`). Worker-compute panes are
7
- # tmux-pane backend siblings. Each wrapper/dispatcher tags the pane it owns with
8
- # a pane-level user option (`@okstra_trace_run=<RUN_DIR>` or
9
- # `@okstra_worker_run=<RUN_DIR>`), so panes are found server-wide by tag — no
10
- # tmux env var or pane-id registry is needed, and the run-scoped tag keeps
11
- # concurrent okstra runs from closing each other's panes.
5
+ # Worker-compute panes are tmux-pane backend siblings. Their dispatcher tags
6
+ # each pane it owns with a pane-level user option (`@okstra_worker_run=<RUN_DIR>`),
7
+ # so panes are found server-wide by tag — no tmux env var or pane-id registry is
8
+ # needed, and the run-scoped tag keeps concurrent okstra runs from closing each
9
+ # other's panes.
10
+ #
11
+ # Trace panes were `tail -F` siblings the provider wrappers split, tagged
12
+ # `@okstra_trace_run` / `@okstra_status`. Those wrappers are now four-line
13
+ # entrypoints and worker progress renders into the worker's own pane, so
14
+ # NOTHING SPAWNS A TRACE PANE and neither tag has a writer left. The trace
15
+ # paths below still run and simply match nothing; `--reclaim-completed`, which
16
+ # keys on `@okstra_status`, is inert for the same reason. Kept rather than
17
+ # deleted because the hooks that call this script are already seeded on user
18
+ # machines — retiring the trace machinery is a deliberate follow-up.
12
19
  #
13
20
  # Two invocation shapes:
14
21
  #
@@ -1,27 +1,34 @@
1
1
  #!/usr/bin/env python3
2
- """okstra-wrapper-status.py — heartbeat sidecar writer for codex/antigravity wrappers.
3
-
4
- The codex/antigravity wrappers (`okstra-codex-exec.sh`, `okstra-antigravity-exec.sh`)
5
- dispatch a long-running CLI under `Bash(run_in_background: true)` and rely on
6
- `BashOutput` polling for liveness. That polling stream only carries stdout
7
- plus a binary `running`/`completed` state. Several recovery decisions need
8
- more — specifically, "did this wrapper start at all, when, and how did it
9
- finish?" — so the wrappers write a small JSON sidecar at
10
- `<prompt-path>.status.json` that survives independent of the polling channel.
11
-
12
- Consumers:
13
-
14
- * `codex-worker` / `antigravity-worker` step 8c: read `log_path` to capture a
15
- diagnostic tail when `exit_code == 0` but the canonical Result file is
16
- absent.
17
- * Lead: cross-check `started_ts` / `ended_ts` to distinguish "wrapper hung
2
+ """okstra-wrapper-status.py — standalone CLI for the worker status sidecar.
3
+
4
+ A CLI worker dispatch runs a long-running CLI in the background, and the
5
+ polling channel that watches it carries only stdout plus a binary
6
+ `running`/`completed` state. Several recovery decisions need more —
7
+ specifically, "did this worker start at all, when, and how did it finish?" — so
8
+ a small JSON sidecar is written at `<prompt-path>.status.json` that survives
9
+ independent of the polling channel.
10
+
11
+ **Nothing calls this script.** Every provider entrypoint now runs through
12
+ `scripts/okstra_ctl/worker_runner.py`, which writes the same document
13
+ in-process; the shell wrappers that shelled out to this file were its only
14
+ callers and they are gone. It is kept for now rather than deleted because
15
+ removing it also means touching the build sync list, the install payload, its
16
+ own test module and the docs that name it — a deliberate follow-up, not a side
17
+ effect of the wrapper migration. `scripts/okstra_ctl/wrapper_status.py` is the
18
+ reader, and what it reads is what the runner writes.
19
+
20
+ Consumers of the sidecar:
21
+
22
+ * A CLI worker's diagnostic step: read `log_path` to capture a tail when
23
+ `exit_code == 0` but the canonical Result file is absent.
24
+ * Lead: cross-check `started_ts` / `ended_ts` to distinguish "worker hung
18
25
  before CLI launched" from "CLI finished but never wrote artifact" when
19
26
  applying the redispatch policy (see team-contract "Lead Redispatch
20
27
  Policy on Result-Missing").
21
28
 
22
- Failures are deliberately non-fatal for the caller — the wrapper's main
23
- job is to run the underlying CLI; a missing sidecar must not break that.
24
- On any error the script prints a one-line diagnostic to stderr and exits 0.
29
+ Failures are deliberately non-fatal for the caller — the caller's main job is
30
+ to run the underlying CLI; a missing sidecar must not break that. On any error
31
+ the script prints a one-line diagnostic to stderr and exits 0.
25
32
 
26
33
  Schema (schemaVersion 1):
27
34
 
@@ -57,7 +57,7 @@ Use the screen to tell "still working" from "stuck", and to see at a glance whic
57
57
  - Worker completion is valid only from `workerDispatches[]`, terminal status sidecars, and required Result Paths. Pane creation alone is not completion.
58
58
  - Reverify uses a fresh jobs file at `runs/<task-type>/state/reverify-jobs-r<N>-<task-type>-<seq>.json`, sets `dispatchKind: "reverify-r<N>"`, and dispatches with `okstra team dispatch --project-root <root> --run-manifest <path> --dispatch-kind reverify-r<N> --jobs-file <jobs-file>`.
59
59
  - Report-writer uses a fresh one-job jobs file with `dispatchKind: "report-writer"` and the same schema, then dispatches through `okstra team dispatch --project-root <root> --run-manifest <path> --jobs-file <jobs-file>`.
60
- - Every reverify or report-writer jobs file carries `workerId`, `provider`, `role`, `modelExecutionValue`, `promptPath`, `resultPath`, `workerResultPath`, and `completionPaths`. For reverify, set `role` to `worker-reverify-r<N>` for the pane title. The report-writer completion paths include both data.json and the worker-results audit file.
60
+ - Every reverify or report-writer jobs file carries `workerId`, `provider`, `role`, `modelExecutionValue`, `promptPath`, `resultPath`, `workerResultPath`, and `completionPaths`. For reverify, set `role` to `worker-reverify-r<N>` — the role selects the dispatch's idle budget and is recorded in the run's status sidecar, so it must name the actual assignment. The report-writer completion paths include both data.json and the worker-results audit file.
61
61
  - After either dispatch, run `okstra team await --project-root <root> --run-manifest <path>` before evaluating terminal status or completion paths.
62
62
 
63
63
  ## Completion, cleanup, and resume
@@ -540,7 +540,7 @@ Schema rules:
540
540
 
541
541
  ## Coverage critic pass
542
542
 
543
- Runs only when `convergence.critic.enabled == true` (set by `--critic <provider>` or the okstra-run `critic_pick` step; default off). Applies to the three finding-producing phases (`requirements-discovery`, `error-analysis`, `implementation-planning`); for `final-verification` the critic runs in a different mode — see §"Acceptance critic pass (final-verification)". This pass targets **coverage** (missed findings), distinct from convergence which targets **agreement quality**.
543
+ Runs only when `convergence.critic.enabled == true` (set by `--critic <provider>` or the okstra-run `critic_pick` step; default off). Applies to the three finding-producing phases (`requirements-discovery`, `error-analysis`, `implementation-planning`); for `final-verification` the critic runs in a different mode — see §"Acceptance critic pass (final-verification)". This pass targets **scope in both directions** — findings that are missing (coverage) and work the findings propose that no requirement asked for (over-scope) — distinct from convergence, which targets **agreement quality** among the findings already raised. The pass keeps its `coverage` mode id and `gaps` vocabulary for both halves; the two are told apart by each candidate's `category`, so no schema or reducer distinguishes them.
544
544
 
545
545
  ### When
546
546
 
@@ -554,10 +554,12 @@ Dispatch one fresh pass to `config.critic.provider` through `redispatch_worker`,
554
554
 
555
555
  The `-worker-` token is load-bearing, not decoration: the critic prompt carries the same generated anchor headers as every other worker ([team-contract](./team-contract.md) §"Worker prompts"), and its `**Audit sidecar path:**` comes from passing that result path through `okstra_ctl.worker_artifact_paths.audit_sidecar_rel()`, which inserts `-audit-` after the token and raises without it. A `<provider>-critic-...` name leaves the lead choosing between breaking the contract and hand-inventing the sidecar name. Note that `originWorker` stays `"<provider>-critic"` — that is a worker id in the convergence state, not a filename, and the two do not have to match.
556
556
 
557
- The critic prompt seeds the consolidated findings and asks ONLY for coverage gaps:
557
+ The critic prompt carries the full analysis contract those anchor headers belong to — worker preamble, error-contract path, audit sidecar, packet boundary — but it is **exempt from the initial-analysis equality group**. Initial analysis workers must receive byte-identical normalized bodies (`worker_prompt_contract.validate_analysis_prompt_set`), and a critic body is deliberately unlike them, so `okstra_ctl.worker_prompt_policy.resolve_prompt_plan` resolves `dispatchKind = "critic"` to `equality_group = None` while keeping `audience = "analysis"`. Never reshape a critic prompt to match the initial one to satisfy that check — matching it would delete the pass. **Enforced:** `tests/contract/test_critic_prompt_equality_exemption.py` pins both directions — the critic is exempt, and two mismatched *initial* prompts still fail.
558
558
 
559
- Required reading before proposing a gap:
560
- - the current run's `analysis-packet.md` for requirements and phase scope;
559
+ The critic prompt seeds the consolidated findings and asks for two things only — coverage gaps and unrequested work. The second half exists because every other scope check in the lifecycle runs *later*: the plan-body `P-Opt-*` verdict judges options this phase has not authored yet (report-writer writes the plan in Phase 6), and the `implementation` verifier judges a diff. Here the unit is a **finding**, which is the input those stages build on — catching unrequested work while it is still a finding is the cheapest place to catch it at all.
560
+
561
+ Required reading before proposing a gap or an over-scope candidate:
562
+ - the current run's `analysis-packet.md` for requirements and phase scope — for the over-scope half this is the authority you search against, so read it before judging any finding unrequested;
561
563
  - `convergence-groups-<task-type>-<seq>.json` for the complete Round 0 ledger;
562
564
  - every initial analysis-worker result named by team-state;
563
565
  - each matching audit sidecar, to distinguish an uninspected path from a claim that was inspected but summarized during grouping.
@@ -565,13 +567,30 @@ Required reading before proposing a gap:
565
567
  Operational guardrails are not task requirements. A gap must trace to a brief requirement, an analysis-packet scope item, a source path the packet authorizes, or an evidence claim in a worker result. Do NOT infer missing verification from a one-line summary; open the named result and audit sidecar first.
566
568
 
567
569
  ```
568
- You are the coverage critic for <task-key>. Below are the consolidated findings
569
- the workers produced. Your ONLY job is to name what is MISSING:
570
+ You are the scope critic for <task-key>. Below are the consolidated findings the
571
+ workers produced. Your job has exactly two halves. Answer both.
572
+
573
+ (1) MISSING — name what nobody covered:
570
574
  - files / directories / execution paths nobody inspected,
571
575
  - requirements or acceptance points with zero findings,
572
576
  - claims raised but never verified.
573
- For each gap, emit a NEW finding with evidence (file:line or the requirement quote).
574
- Do NOT restate an existing finding. If nothing is missing, say so explicitly.
577
+ For each, emit a NEW finding with evidence (file:line or the requirement quote).
578
+
579
+ (2) UNREQUESTED — name work these findings propose that no requirement asked for:
580
+ - a finding whose proposed change serves no requirement, scope item, or
581
+ acceptance point you can QUOTE from the analysis packet,
582
+ - an abstraction, configuration knob, or generalization proposed for a caller or
583
+ a case nobody has stated,
584
+ - a rewrite, migration, or cleanup of code the requirements never mention.
585
+ For each, emit a candidate with `category: "unrequested-scope"`, quote the
586
+ proposed work verbatim, and state which requirement you searched for and did not
587
+ find.
588
+
589
+ Do NOT restate an existing finding. Judge (2) against the analysis packet's
590
+ requirements and scope, never against your own preference for how the code should
591
+ look — "I would have done it differently" is not unrequested work, and neither is
592
+ work the packet authorizes but you consider unnecessary. If a half has nothing,
593
+ say so explicitly for that half; silence on one half is an incomplete result.
575
594
  ```
576
595
 
577
596
  ### Gap verification (1 adversarial reverify round)
@@ -579,8 +598,17 @@ Each critic gap enters the verification queue as a finding with `originWorker =
579
598
 
580
599
  **A gap that received no verdict is NOT a rejected gap (BLOCKING).** Dropping applies only to gaps the voters actually judged. A gap can also end the round *unjudged* — the verification dispatch returned a terminal non-result (`timeout`, `error`, no result file), the returned result covered only some of the gaps, or no non-critic analyser was available to vote at all. Nobody inspected those, so classifying them as hallucinations is a fabricated verdict. Each one MUST be recorded as a `## 5. Missing Information and Risks` row (`missingInformation`, `source: "critic-unverified"`) whose `risk` names the gap and the reason verification did not complete, and counted in `config.critic.gapsUnverified`. They are **not** promoted to findings (unverified) and **not** raised as `clarification` items — an unverified gap needs an analyser to verify it on the next run, not a decision from the user. Silently losing them is a contract violation: the batch that times out is exactly the batch of gaps too expensive to check, so the highest-risk items are the ones that vanish.
581
600
 
601
+ **`category: "unrequested-scope"` candidates are classified the same way but disposed of differently.** A coverage gap the voters contest is a hallucination — nothing was actually missing, so dropping it costs one wasted verification. An over-scope candidate the voters contest is a *disagreement about whether the work was asked for*, and dropping that silently returns the run to the state this half exists to change. So:
602
+
603
+ - `full-consensus` / `partial-consensus` → merge as a finding, exactly like a coverage gap. The merged finding names the unrequested work and the requirement search that came up empty; the phase's own deliverable rules decide whether it lands as a dropped item or a clarification row.
604
+ - `contested` / `worker-unique` → counted in `config.critic.gapsRejected` (unchanged accounting) but ALSO recorded as a `## 5. Missing Information and Risks` row with `source: "critic-unconfirmed-scope"`, whose `risk` quotes the proposed work and the split verdict. It is **not** promoted to a finding: a contested over-scope claim must not block a plan on one worker's taste. The user reads the row and decides.
605
+ - no verdict at all → the `critic-unverified` rule above applies unchanged.
606
+
607
+ The asymmetry is deliberate and runs the opposite way from the coverage half: a false "you missed something" costs a verification, while a false "you built too much" costs a real requirement — so the first may be dropped outright and the second is recorded either way. This is the same reasoning that makes the `final-verification` acceptance critic never drop a candidate, applied to the one direction where dropping is otherwise the default.
608
+
582
609
  ### State
583
610
  - `convergence.critic` manifest block: `{ enabled, provider, modelExecutionValue }`.
611
+ - Each candidate's `category` tells the two halves apart: literal `"unrequested-scope"` for the over-scope half, any other value for a coverage gap. `schemas/convergence-critic-results-v1.0.schema.json` leaves `category` a free string, so this needs no schema or reducer change — but it also means nothing machine-checks the spelling. A misspelled category is read as a coverage gap and silently takes the drop-on-contested path.
584
612
  - The lead passes one canonical coverage batch with `{ schemaVersion, taskKey, mode, provider, modelExecutionValue, dispatches, gaps }`; each gap carries its candidate fields plus `gapId` and `votes`. `dispatches[]` contains exactly one row for every Phase 4 analyser except the critic, even when execution did not produce a result: persist `status: timeout | error | not-run` and the elapsed `durationMs` instead of omitting that analyser. `apply-critic-gaps` rejects a non-terminal main queue, duplicate or missing analysers, unknown workers, critic dispatches/votes, votes without a completed dispatch, and a second batch.
585
613
  - Convergence state artifact: merged gaps appear in `findings[]` with `source: "critic"` and `rounds: []`. The separate `criticVerification.gaps[]` ledger retains each gap's `summary`, `category`, `ticketIds`, `originEvidence`, optional `evidenceArtifacts`, classification, merge link, and votes. Strict v1.3 validation deterministically replays each complete critic-origin finding from that ledger; a critic batch never increments `roundHistory` or `totalRounds` and never creates a fake main round.
586
614
  - `config.critic` is `{ provider, modelExecutionValue, gapsProposed, gapsMerged, gapsRejected, gapsUnverified }`, with `gapsProposed = gapsMerged + gapsRejected + gapsUnverified`. `full-consensus` / `partial-consensus` gaps merge, `contested` / `worker-unique` gaps count as rejected, and gaps with no usable analyser vote appear in both the ledger and final `unverifiedGaps[]`.
@@ -104,9 +104,10 @@ Required checkpoints:
104
104
  - `PROGRESS: phase-5-collect worker=<role> status=<terminal-status>` — once per worker, immediately after the result file is verified.
105
105
  - `PROGRESS: phase-5.5-convergence round=<N> queue=<count>` — at the start of each convergence round (Phase 5.5).
106
106
  - `PROGRESS: phase-5.6-critic provider=<provider> gaps=<n>` — after the critic result is collected (Phase 5.6, opt-in; the critic dispatch itself fires concurrently with the first 5.5 reverify round). Omitted when `convergence.critic.enabled == false`.
107
- - `PROGRESS: phase-batch-cleanup panes=<n>` — immediately after cleaning up the previous batch's panes, at each batch boundary (① just before the first `phase-5.5-convergence` round ② just before the `phase-6-synthesis` report-writer dispatch). `<n>` is the number of panes reclaimed at that boundary — trace panes plus completed teammate panes, which are panes too — read from the cleanup's `--list` pass taken immediately before the reclaim, never estimated. Expose only the counts and NEVER expose `%NNN`/lead-pane.id/raw worker handles. Just before the first batch (analysis-worker dispatch) there is nothing to clean up, so it is a no-op and the marker is omitted.
107
+ - `PROGRESS: phase-batch-cleanup panes=<n>` — immediately after cleaning up the previous batch's panes, at each batch boundary (① just before the first `phase-5.5-convergence` round ② just before the `phase-6-synthesis` report-writer dispatch). `<n>` is the number of panes reclaimed at that boundary — worker-compute panes plus completed teammate panes, which are panes too — read from the cleanup's `--list` pass taken immediately before the reclaim, never estimated. Expose only the counts and NEVER expose `%NNN`/lead-pane.id/raw worker handles. Just before the first batch (analysis-worker dispatch) there is nothing to clean up, so it is a no-op and the marker is omitted.
108
108
  - `PROGRESS: phase-6-synthesis dispatching report-writer-worker` — at the start of Phase 6.
109
109
  - `PROGRESS: phase-5.5.9-plan-verify round=<N> items=<count>` — immediately before dispatching each plan-body verification round (`implementation-planning` only; see [plan-body-verification](./plan-body-verification.md) §"Round protocol"). Each round is a worker batch like any other, so round 2 and later MUST be preceded by a `phase-batch-cleanup` line reclaiming the previous round's verifiers. The numbering keeps this line sorted where the work happens — after Phase 6, because the round verifies the drafted plan body.
110
+ - `PROGRESS: user-confirm <C-NNN> <the question, one line>` — immediately before asking the user about anything that would otherwise become an open `Blocks=approval` row (see "User confirmation before an approval blocker" below). Not tied to a phase: it fires wherever the blocker surfaces. `<C-NNN>` is the id the row will carry, so the answer and the row can be matched afterwards.
110
111
  - `PROGRESS: phase-7-persist updating manifests` — at the start of Phase 7.
111
112
  - `PROGRESS: phase-7-teardown shutting-down-workers` — only after usage collection and user approval, immediately before `shutdown_workers`; omitted when no cleanup resource exists or the user keeps it.
112
113
  - `PROGRESS: complete final-report=<relative-path>` — final summary line, after all persistence.
@@ -117,6 +118,27 @@ Do NOT replace them with prose ("Now I'm starting Phase 2..."), do NOT skip a ch
117
118
 
118
119
  **Enforcement:** the Phase 7 validator (`validators/validate-run.py` → `validate_session_conformance.py`) reads the selected adapter's declared conformance evidence/event source within the run window and fails the run as `contract-violated` when a required checkpoint is missing — including the per-worker `phase-4-dispatch` / `phase-5-collect` lines (which must name each worker's role) and the `phase-batch-cleanup` lines that MUST precede the first `phase-5.5-convergence` round and the `phase-6-synthesis` report-writer dispatch. When the plan-body state file records two or more rounds, `_check_plan_verify_cleanup_checkpoints` additionally requires a `phase-5.5.9-plan-verify` line per round and a `phase-batch-cleanup` between consecutive rounds — that boundary sat outside both older checks, so a five-round self-fix loop left every round's verifiers holding their panes. `phase-7-teardown` and `complete` fire after validation and are not checked.
119
120
 
121
+ ## User confirmation before an approval blocker (BLOCKING)
122
+
123
+ An open `Blocks=approval` row stops the whole task: the plan cannot be approved, `implementation` cannot start, and the answer arrives only through a separate user-response cycle and a re-run. It is the most expensive artifact this contract lets the lead produce. **Before writing one, ask the user.**
124
+
125
+ This is not a phase. It fires wherever the blocker surfaces — during intake when the directive and the brief disagree, mid-convergence when workers split on something only the user can settle, in the §5.5.9 self-fix loop when an item no round can clear keeps the gate red.
126
+
127
+ The sequence is fixed:
128
+
129
+ 1. Emit `PROGRESS: user-confirm <C-NNN> <the question, one line>` with the id the row would carry.
130
+ 2. Ask in plain user-facing text: what is undecided, the options with their consequences, and which one you recommend. One question at a time.
131
+ 3. On an answer — record it in the row's `userInput`, set `status: answered` and `userConfirmation: asked-and-answered`, apply it, and **keep going in this run**. An answered question is not a blocker, and a run that stops anyway wastes the answer it just received.
132
+ 4. Only when asking fails does the row stay open: `asked-awaiting` when the user has not answered, `deferred-no-interactive-session` when this run has no user to ask.
133
+
134
+ **Predicting the blocker is not the same as raising it.** A lead that says "this will likely become an approval blocker; I will ask at that point" has already reached the moment — ask then, in that message. One run announced exactly that, never asked, wrote the row anyway, and then spent its entire self-fix budget on a gate no round could clear, because the user had already answered the question before the run started.
135
+
136
+ **`lead-directed` blockers cannot be deferred.** When the item is the lead's own judgment rather than a worker's finding, and this run has nobody to ask, the row is not the outlet — record a Working Assumption in `## 5. Missing Information and Risks` naming the assumption the plan proceeds under, exactly as a surviving planner-fixable item does, and let the plan proceed. Blocking a plan on the lead's own judgment in a run where that judgment cannot be put to the user only moves the work to a re-run.
137
+
138
+ **Never seed the answer into a worker prompt.** Instructing a worker to "raise this as a user decision rather than choosing" and then reporting the resulting agreement as an independent finding misrepresents where the blocker came from. If it is the lead's judgment, `origin` is `lead-directed` — see [_common-contract.md](../profiles/_common-contract.md) "Clarification request policy".
139
+
140
+ **Enforcement:** `validators/validate-run.py` `_validate_open_approval_blocker_provenance` fails any open approval blocker missing `origin` / `userConfirmation`, and any `lead-directed` one deferred for want of an interactive session; `validators/validate_session_conformance.py` `_check_user_confirm_checkpoints` fails a row claiming the user was asked when no matching `user-confirm` line exists in this run's evidence.
141
+
120
142
  ## Model assignments
121
143
 
122
144
  **The lead never invents a model.** Every role's model is read from `task-manifest.json` → `resultContract.requiredWorkerRoles[*].modelExecutionValue` (and the lead model metadata). A missing assignment is a manifest defect, not a license to fall back — see [team-contract](./team-contract.md) "Model Assignment Rules". The manifest is always populated at run-prep time by the CLI, which seeds these values from `OKSTRA_DEFAULT_*_MODEL` (`scripts/okstra_ctl/run.py`).