okstra 0.209.5 → 0.209.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/package.json +1 -1
- package/runtime/BUILD.json +2 -2
- package/runtime/python/okstra_ctl/phases/implementation/instructions/_implementation-executor.md +1 -1
- package/runtime/python/okstra_ctl/phases/implementation_planning/profile.md +1 -1
- package/runtime/python/okstra_ctl/write_policy.py +10 -3
package/package.json
CHANGED
package/runtime/BUILD.json
CHANGED
package/runtime/python/okstra_ctl/phases/implementation/instructions/_implementation-executor.md
CHANGED
|
@@ -89,7 +89,7 @@ template's check; that template is gone.
|
|
|
89
89
|
```
|
|
90
90
|
|
|
91
91
|
The file MUST NOT exist before the run starts (overwrite is refused — see `--force-stage` non-goal). **Enforced:** `validators/validate-run.py` `_validate_stage_carry_sidecar_exists` fails a run that declares `stageSidecarEvidence` without the file on disk. Transcribing the JSON into the report is not the same as writing it: `consumers` treats the carry file as the source of truth for marking the stage `done`, so a missing file leaves the stage permanently incomplete and blocks every dependent stage with a `PrepareError` — while this run reports success.
|
|
92
|
-
- **Verifier gates are not yours to run (BLOCKING).** The self-mock detector (`validators/detect_self_mock.py`) belongs to the implementation verifier and is never delegated to you (`_implementation-verifier.md` §"Self-mock detection"). The coding-conventions preflight names it as the enforcement behind the no-self-mocking principle — that names who will check your diff, not a command for you to run. You MUST NOT invoke it and MUST NOT write `<task_root>/qa/self-mock-*.json`; running it early does not pre-satisfy the gate, because the verifier runs it again under its own duty. **Enforced:** that sidecar is not among the paths your attempt's `writePolicy.artifactPolicy.allowedPaths` carries, so the write audit closes the attempt as `error` with `artifact-root change exceeds batch policy union` (`scripts/okstra_ctl/execution_mutation_audit.py`) — a stage whose every gate passed still lands as a failed run. The same holds for the conformance results `<task_root>/qa/result-*.json`: you may run a conformance script to check your work —
|
|
92
|
+
- **Verifier gates are not yours to run (BLOCKING).** The self-mock detector (`validators/detect_self_mock.py`) belongs to the implementation verifier and is never delegated to you (`_implementation-verifier.md` §"Self-mock detection"). The coding-conventions preflight names it as the enforcement behind the no-self-mocking principle — that names who will check your diff, not a command for you to run. You MUST NOT invoke it and MUST NOT write `<task_root>/qa/self-mock-*.json`; running it early does not pre-satisfy the gate, because the verifier runs it again under its own duty. **Enforced:** that sidecar is not among the paths your attempt's `writePolicy.artifactPolicy.allowedPaths` carries, so the write audit closes the attempt as `error` with `artifact-root change exceeds batch policy union` (`scripts/okstra_ctl/execution_mutation_audit.py`) — a stage whose every gate passed still lands as a failed run. The same holds for the conformance results `<task_root>/qa/result-*.json`: you may run a conformance script to check your work — a baseline capture or `--compare` target goes under `<task_root>/qa/baseline/` and anything else it writes (build logs, measurements) under `<task_root>/qa/output/`, the only qa directories besides `qa/scripts/` your attempt may write — but the verifier records every result, including an earlier stage's result that your rewrite of its script made stale. When the approved plan tells you to write one, leave it and name it in your result as a plan step the verifier owns.
|
|
93
93
|
- **An external Tier 3 non-PASS does NOT withhold the carry evidence.** A Tier 3 entry whose `requires` include `http`, `external`, or `db` is advisory. Its FAIL, MISSING, no result, startup failure, or credential / network / service absence gets recorded honestly — exact command, exit code, output tail, marked `ADVISORY` in `Validation evidence` — and you emit the carry evidence anyway. Only Tier 1 and Tier 2 failures withhold it. Withholding on an external result is what actually blocks the stage: the carry file is the only thing that can mark a stage `done`, the verifier re-runs that same command from the host (where a call your sandbox could not complete often passes), and a stage the verifier then PASSes can never be closed because its evidence was never written.
|
|
94
94
|
- **Reverse link (BLOCKING).** The runtime already appended a `status:"started"` row for this stage before the run began. The terminal row belongs to the lead's post-stage persistence and is verdict-gated — `status:"done"` with `carry_path` on a non-`FAIL` verdict, `status:"failed"` on `FAIL` (`_implementation-deliverable.md` §"Lead post-stage persistence").
|
|
95
95
|
- **No PR / push in this phase.** This run produces local commits, carry sidecar evidence, verifier results, and the implementation final report only. Push and PR creation belong exclusively to the later `release-handoff` phase after `final-verification` returns `accepted`.
|
|
@@ -160,7 +160,7 @@ Plan for the actual worktree layout before approval. Shared documentation direct
|
|
|
160
160
|
unavailable environment is a user-owned follow-up, never a plan approval or
|
|
161
161
|
later run blocker. `requires=[]` and `requires=[io]` remain blocking.
|
|
162
162
|
Remote IO should also declare `external`.
|
|
163
|
-
Layout split (not this phase's writes): executable scripts (conformance + any real-IO test) live under `<task_root>/qa/scripts/`; data sidecars (`conformance-manifest.json`, `result-*.json`) stay at the `qa/` root. A step that
|
|
163
|
+
Layout split (not this phase's writes): executable scripts (conformance + any real-IO test) live under `<task_root>/qa/scripts/`; data sidecars (`conformance-manifest.json`, `result-*.json`) stay at the `qa/` root. A step that captures a baseline or a `--compare` target points it at `<task_root>/qa/baseline/`; anything else a script writes on every run (build logs, measurements, captured responses) goes under `<task_root>/qa/output/`. These are the only other qa directories the implementer's write audit accepts, and any other qa path discards the attempt (`scripts/okstra_ctl/write_policy.py` `_role_qa_artifact_paths`). The split matters at final verification: the verifier re-runs every Tier 3 script and may rewrite `qa/output/`, but `qa/baseline/` must stay byte-identical (`_inherited_qa_evidence`), so a baseline written under `qa/output/` is not protected. The implementer writes the scripts and the manifest entry. Only the implementation verifier writes `result-*.json`, including the result of an earlier stage's entry whose script this stage rewrites, so a step never assigns a result file to the implementer and never lists one in `plannedPaths`. This declaration is enforced at four layers: `validators/validate-implementation-plan-stages.py` check **S11** forces every stage to carry one of the two lines; at the planning boundary `scripts/okstra_ctl/phases/implementation_planning/plan_body.py` `_validate_planning_conformance_declared` accepts a well-formed `Conformance tests:` line even when the script file and manifest entry are absent (malformed `requires` still fails); the matching `implementation` stage run that inherited `Conformance tests:` fails closed when the script file is missing (`_validate_conformance`); and the manifest JSON structure — including each entry's `script` living under `qa/scripts/` and a `runCommand` that does not change cwd — is enforced by `validate_conformance_manifest` when the implementer writes the entry.
|
|
164
164
|
- `### Stage Exit Contract` — predicted added/modified files, newly exposed identifiers/types/endpoints, downstream-usable resources.
|
|
165
165
|
- `### Stage Validation` — pre / mid / post exact commands or observable outcomes for this stage only.
|
|
166
166
|
- **Run-executable only (BLOCKING).** A `validationChecklist` row that carries `stageRefs` gates that stage's carry, so it MUST be executable by the implementation run itself — no deployment, no manual walk-through, no observation "by hand", no credentials the run does not hold. The implementation phase forbids deploys and holds no deploy credentials, so such a row can never pass and blocks a stage with zero code defects (observed: dev-10341 VC-013/VC-014). Manual or deployed-environment verification belongs in the brief's `External Gates` and final-verification's user-owned external QA — record it there without `stageRefs`. **Enforced:** `okstra_ctl.implementation_direction.validate_selected_direction_plan` (`stage_validation_executability_errors`) rejects stage-gating rows whose observation text requires a manual or deployed-environment step.
|
|
@@ -101,8 +101,10 @@ def _role_qa_artifact_paths(role: str, task_root: Path | None) -> tuple[Path, ..
|
|
|
101
101
|
`qa/scripts/`, plus the manifest entry naming them
|
|
102
102
|
(`_implementation-executor.md` §"Stage conformance script", §"Real-IO test
|
|
103
103
|
isolation"), plus `qa/output/` for whatever those scripts write when the
|
|
104
|
-
executor runs them against its own work —
|
|
105
|
-
(§"Verifier gates are not yours to run" lets it run them)
|
|
104
|
+
executor runs them against its own work — build logs, measurements
|
|
105
|
+
(§"Verifier gates are not yours to run" lets it run them) — and
|
|
106
|
+
`qa/baseline/` for a captured baseline or `--compare` target, which the
|
|
107
|
+
final verifier must not rewrite. Everything else
|
|
106
108
|
under `qa/` — the self-mock sidecar and its diff, the conformance run's
|
|
107
109
|
`result-*.json` — is written while the verifier runs its own gates.
|
|
108
110
|
Keeping the executor's grant to those three entries is
|
|
@@ -124,6 +126,7 @@ def _role_qa_artifact_paths(role: str, task_root: Path | None) -> tuple[Path, ..
|
|
|
124
126
|
return (
|
|
125
127
|
task_qa_dir(task_root) / "scripts",
|
|
126
128
|
task_qa_dir(task_root) / "output",
|
|
129
|
+
task_qa_dir(task_root) / "baseline",
|
|
127
130
|
task_conformance_manifest_file(task_root),
|
|
128
131
|
)
|
|
129
132
|
if role == "verifier":
|
|
@@ -154,7 +157,9 @@ def _inherited_qa_evidence(
|
|
|
154
157
|
baseline overwrote the implementation's RED evidence and still passed the
|
|
155
158
|
audit (observed 2026-09-26, dev-11054 `qa/baseline/base.json`). New files and
|
|
156
159
|
the conformance results it re-runs (`qa/result-<stageKey>.json`, the only qa
|
|
157
|
-
write `final_verification/boundary.json` names) stay writable
|
|
160
|
+
write `final_verification/boundary.json` names) stay writable, and so does
|
|
161
|
+
`qa/output/`: the Tier 3 scripts it must re-run rewrite their own outputs
|
|
162
|
+
there (dev-11127 final-verification); baselines live in `qa/baseline/`.
|
|
158
163
|
"""
|
|
159
164
|
if role != "verifier" or task_type != "final-verification" or task_root is None:
|
|
160
165
|
return ()
|
|
@@ -165,9 +170,11 @@ def _inherited_qa_evidence(
|
|
|
165
170
|
conformance_result_file(qa_dir, key)
|
|
166
171
|
for key in _conformance_stage_keys(task_root)
|
|
167
172
|
}
|
|
173
|
+
regenerated = qa_dir / "output"
|
|
168
174
|
return tuple(sorted(
|
|
169
175
|
path for path in qa_dir.rglob("*")
|
|
170
176
|
if path.is_file() and not path.is_symlink() and path not in rerun
|
|
177
|
+
and not path.is_relative_to(regenerated)
|
|
171
178
|
))
|
|
172
179
|
|
|
173
180
|
|