okstra 0.204.0 → 0.205.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +4 -3
- package/docs/architecture.md +1 -1
- package/docs/performance-improvement-plan-v2.md +1 -6
- package/docs/project-structure-overview.md +4 -2
- package/docs/task-process/implementation-option-selection.md +1 -1
- package/docs/task-process/release-handoff.md +2 -2
- package/package.json +1 -1
- package/runtime/BUILD.json +2 -2
- package/runtime/prompts/lead/okstra-lead-contract.md +2 -2
- package/runtime/prompts/profiles/_implementation-diff-review.md +1 -1
- package/runtime/prompts/profiles/_implementation-self-check.md +1 -1
- package/runtime/prompts/profiles/final-verification.md +1 -1
- package/runtime/prompts/profiles/release-handoff.md +8 -6
- package/runtime/prompts/wizard/prompts.ko.json +3 -2
- package/runtime/python/okstra_ctl/consumers.py +17 -5
- package/runtime/python/okstra_ctl/convergence_engine.py +38 -0
- package/runtime/python/okstra_ctl/handoff.py +221 -13
- package/runtime/python/okstra_ctl/handoff_verification.py +25 -6
- package/runtime/python/okstra_ctl/incremental_carry.py +19 -2
- package/runtime/python/okstra_ctl/plan_items_cli.py +9 -0
- package/runtime/python/okstra_ctl/render.py +29 -0
- package/runtime/python/okstra_ctl/wizard/steps_plan.py +11 -0
- package/runtime/validators/validate-run.py +50 -8
- package/runtime/validators/validate_session_conformance.py +70 -1
package/README.md
CHANGED
|
@@ -94,7 +94,7 @@ Development and CI use Node.js 24 LTS and Python 3.14, selected by `.nvmrc` and
|
|
|
94
94
|
├── prompts/ lead operating contracts (prompts/lead/*) + coding-preflight resource pack
|
|
95
95
|
├── installed-runtimes.json installed runtime adapter manifest
|
|
96
96
|
├── installed-skills.json installed skill manifest (used by uninstall)
|
|
97
|
-
├── installed-agents.json
|
|
97
|
+
├── installed-agents.json reclaim record of host agent files a past version installed
|
|
98
98
|
├── recent.jsonl, active.jsonl run indexes
|
|
99
99
|
├── memory-book/ global conversation memory (`okstra memory`)
|
|
100
100
|
├── projects/ per-project metadata mirrors
|
|
@@ -108,8 +108,9 @@ Development and CI use Node.js 24 LTS and Python 3.14, selected by `.nvmrc` and
|
|
|
108
108
|
└── okstra-*/SKILL.md 8 user-facing skills only (setup/brief/run/memory/inspect/schedule/container/manager)
|
|
109
109
|
|
|
110
110
|
~/.okstra/agents/ role, common and operation contracts (okstra reads them; no host discovers them)
|
|
111
|
-
|
|
112
|
-
|
|
111
|
+
├── common.json rules every okstra-owned invocation carries
|
|
112
|
+
├── roles/*.json one contract per role (analyser, critic, verifier, …)
|
|
113
|
+
└── operations/*.json work that runs outside a lifecycle run (code review, translation, …)
|
|
113
114
|
|
|
114
115
|
<project-root>/.okstra/
|
|
115
116
|
├── project.json {projectId, projectRoot, ...} (written by `/okstra-setup`)
|
package/docs/architecture.md
CHANGED
|
@@ -997,7 +997,7 @@ Fragments are substrings, so an over-broad fragment silently promotes neighbouri
|
|
|
997
997
|
- Worktree targeting is runner-based. A native-session executor uses its host adapter's edit and command primitives against `EXECUTOR_WORKTREE_PATH`; a CLI-backed executor receives the worktree path through the deterministic provider process. Prefer explicit working-directory flags such as `git -C` or `cargo --manifest-path` when available.
|
|
998
998
|
- The project-level current-task convenience pointer is `.okstra/discovery/latest-task.json`.
|
|
999
999
|
- The project-level canonical task inventory is `.okstra/discovery/task-catalog.json`.
|
|
1000
|
-
- At `okstra install` time, okstra skill assets are seeded to `~/.agents/skills/` by default
|
|
1000
|
+
- At `okstra install` time, okstra skill assets are seeded to `~/.agents/skills/` by default, and to `~/.claude/skills/` as well when `~/.claude` exists (per-project seeding is no longer performed). Agent contracts are not seeded to a host-global agent path: they install to `~/.okstra/agents/` only (ADR-0017).
|
|
1001
1001
|
- Seeded okstra Claude assets instruct Claude to dispatch workers into the session's implicit team with `Agent(name: ...)` (v2.1.178 removed `TeamCreate`/`TeamDelete`). Agent targets receive only skill Markdown.
|
|
1002
1002
|
- The host-native Okstra lead, not preparation scripts or a report writer, makes the final judgment.
|
|
1003
1003
|
- A stable task key must be maintained to enable later bug tracking, corrections, and reverification.
|
|
@@ -89,12 +89,7 @@ This serial rendering has room for improvement, but it is generally cheaper than
|
|
|
89
89
|
|
|
90
90
|
### 2.4 Worker Structure
|
|
91
91
|
|
|
92
|
-
The
|
|
93
|
-
|
|
94
|
-
- `claude-worker.md`: Claude subagent.
|
|
95
|
-
- `codex-worker.md`: invokes the Codex CLI through the `okstra-codex-exec.sh` wrapper.
|
|
96
|
-
- `antigravity-worker.md`: invokes the Antigravity CLI through the `okstra-antigravity-exec.sh` wrapper.
|
|
97
|
-
- `report-writer-worker.md`: final-report author. It is not an analysis worker and does not vote in convergence.
|
|
92
|
+
Superseded. The per-provider worker definition files this section described (`agents/workers/*.md`) were deleted in `ef0d594`; a dispatch now composes the common contract (`agents/common.json`), the role contract (`agents/roles/<role>.json`) and the duty contract (`prompts/duties/<duty>.json`), and the provider is an execution axis rather than a definition file. The measurement below is unaffected: it counts dispatches, not definitions.
|
|
98
93
|
|
|
99
94
|
## 3. Performance Bottleneck Hypotheses
|
|
100
95
|
|
|
@@ -379,6 +379,8 @@ Important modules:
|
|
|
379
379
|
| `domain/worker_stream.py` | the normalised event vocabulary (`Text` / `ToolCall` / `ToolResult` / `Denial` / `Result`) plus its three pure projections: `format_live` (one readable row per event, for the pane), `format_log` (the same plus bodies, for the archive), `final_text` (the closing message alone). Also `content_block_events`, the normaliser for the wire shape keyed on `type` with `message.content` blocks, which three providers share. No files, no clock |
|
|
380
380
|
| `domain/worker_role.py` | per-role execution budgets — the 1500s/600s idle pair lives here once instead of being re-declared in each wrapper |
|
|
381
381
|
| `task_target.py` | shared helper resolving `task-key → (task_root, project_root)` (`resolve_task_root`) |
|
|
382
|
+
| `contract_refreeze.py` | re-freezes a running run's frozen contracts in the installed format. A run freezes its duty/role/common contracts at start and pins their digest, so installing a release that changed the contract format leaves that run unable to produce another prompt (`run duty snapshot catalog digest does not match`). Backs `okstra agent-prompt refreeze-contracts`, rewrites the frozen copies and the manifest digest only, and never touches a prompt, result or ledger |
|
|
383
|
+
| `operation_invocation.py` | prepares an Okstra-owned LLM operation that runs outside `okstra-run` — the operation contract (`agents/operations/<id>.json`) owns the duty and the worker count, and the role is derived from that duty's `roleId`, so a skill passes the operation name instead of assembling role/provider/model itself (ADR-0017). Backs `okstra agent-prompt resolve-operation` |
|
|
382
384
|
| `contract_graph.py`, `contract_graph_cli.py` | runtime-contract graph loader + cross-reference/dependency-closure validator and its `okstra contract-check --root <dir> (--profile\|--operation)` CLI boundary. Loads the agent contract schemas (`common`/`role`/`duty`/`profile`/`operation`), validates known role capabilities, and reports the dependency closure with per-file `path`/`schemaVersion`/`sha256`; an invalid contract raises `ContractGraphError` |
|
|
383
385
|
| `json_boundary.py` | strict JSON persistence boundaries for okstra-owned artifacts — a sealed `ExternalJsonSource` (validated producer + path) is the only way owned JSON is read, and `JsonBoundaryError` names artifact / reason / path when a write cannot satisfy its contract; the SSOT that keeps the model out of internal JSON key/path authorship |
|
|
384
386
|
| `fixed_text.py` | shared scalar-line format for the model-facing fixed-text projections — `scalar` neutralises complex values and control characters (backticks escaped for the code span `line` wraps it in), `block` projects a prose body outside any code span with its backticks intact, `line` renders one static-labelled Markdown list row, and `value_lines` losslessly flattens a JSON-shaped value into fixed name/order/value rows |
|
|
@@ -435,7 +437,7 @@ Token/cost accounting:
|
|
|
435
437
|
| Path | Role |
|
|
436
438
|
|---|---|
|
|
437
439
|
| `launch.template.md` | Lead prompt template rendered for each run |
|
|
438
|
-
| `duties
|
|
440
|
+
| `duties/<duty>.json` | Canonical functional duty contracts composed into every Okstra-owned LLM invocation, alongside the common contract at `agents/common.json` and the role contract at `agents/roles/<role>.json`; `direction-selection-worker` owns direction comparison/validation while `planning-worker` realizes the selected direction; provider/model identity does not select the duty |
|
|
439
441
|
| `profiles/_common-contract.md` | Shared phase contract |
|
|
440
442
|
| `profiles/<task-type>.md` | Phase profiles (single language — runtime always loads from `profiles/`, never a translated mirror) |
|
|
441
443
|
| `implementation-option-selection.md` | Read-only lifecycle profile for candidate comparison or preselected-direction validation before detailed planning |
|
|
@@ -639,7 +641,7 @@ Both Markdown and HTML are derived, not authoring sources. The schema is the con
|
|
|
639
641
|
└── .locks/
|
|
640
642
|
|
|
641
643
|
~/.claude/skills/okstra-*/SKILL.md
|
|
642
|
-
~/.
|
|
644
|
+
~/.okstra/agents/{common.json,roles/*.json,operations/*.json}
|
|
643
645
|
|
|
644
646
|
<PROJECT_ROOT>/.okstra/
|
|
645
647
|
├── project.json
|
|
@@ -63,7 +63,7 @@ This phase does not edit source code, run builds or tests, execute migrations, d
|
|
|
63
63
|
## 7. Verified code
|
|
64
64
|
|
|
65
65
|
- [`prompts/profiles/implementation-option-selection.md`](../../prompts/profiles/implementation-option-selection.md)
|
|
66
|
-
- [`prompts/duties/direction-selection-worker.
|
|
66
|
+
- [`prompts/duties/direction-selection-worker.json`](../../prompts/duties/direction-selection-worker.json)
|
|
67
67
|
- [`scripts/okstra_ctl/implementation_options.py`](../../scripts/okstra_ctl/implementation_options.py)
|
|
68
68
|
- [`scripts/okstra_ctl/implementation_direction.py`](../../scripts/okstra_ctl/implementation_direction.py)
|
|
69
69
|
- [`scripts/okstra_ctl/exact_coverage.py`](../../scripts/okstra_ctl/exact_coverage.py)
|
|
@@ -41,7 +41,7 @@ flowchart TD
|
|
|
41
41
|
Confirm --> Render[render-bundle]
|
|
42
42
|
```
|
|
43
43
|
|
|
44
|
-
`release-handoff` has no analysis-worker dispatch. Launch selection still shows any applicable role-count / role-model steps; current-session lead is this session and is listed on the confirmation summary. There is no provider roster multi-pick and no `Use defaults / Customize` fork. Dynamic verifiers are not chosen at launch. `--workers` is not a launch picker, and the runtime forces the worker list to empty. The wizard outcome's `renderArgs` includes `pr-template-path` only for release-handoff. Scope selection finishes before prepare, and the project/global save runs before `render-bundle` via the `config.set pr-template-path` action of `outcome.persistActions[]`. Only the stages that were marked `verified` by
|
|
44
|
+
`release-handoff` has no analysis-worker dispatch. Launch selection still shows any applicable role-count / role-model steps; current-session lead is this session and is listed on the confirmation summary. There is no provider roster multi-pick and no `Use defaults / Customize` fork. Dynamic verifiers are not chosen at launch. `--workers` is not a launch picker, and the runtime forces the worker list to empty. The wizard outcome's `renderArgs` includes `pr-template-path` only for release-handoff. Scope selection finishes before prepare, and the project/global save runs before `render-bundle` via the `config.set pr-template-path` action of `outcome.persistActions[]`. Only the stages that were marked `verified` by a release-ready verification in the Stage Lifecycle Snapshot and are not yet covered by a `pr` become candidates; leaving `--stages` empty takes all of them, and the wizard picker offers an all-stages option beside the individual ones. Both verification scopes mark a stage: a single-stage report marks its own stage, and a whole-task report marks every stage in its `stageReports` whose completed commit the verified head contains.
|
|
45
45
|
|
|
46
46
|
Note that this phase is also a target of task worktree provisioning. The normal flow reuses the implementation/final-verification result of the same task-key. Starting a new task may create a new branch, and it is likely to be blocked at the entry gate's "implementation commit exists" condition.
|
|
47
47
|
|
|
@@ -87,7 +87,7 @@ flowchart TD
|
|
|
87
87
|
Before asking the user whether to push/PR, the lead confirms the following.
|
|
88
88
|
|
|
89
89
|
- The `## Source Verification Report` of the input document (`release-handoff-input.md`) generated by prepare lists the selected stages (`HANDOFF_STAGES`) and the cited report table. The brief is the input of the entry phase, so it does not exist in release-handoff — the user's stage selection finishes before prepare via the wizard `handoff_stage_pick` or the CLI `--stages` (empty takes every eligible stage).
|
|
90
|
-
- Each cited
|
|
90
|
+
- Each cited report must be the latest validated execution for the run it belongs to — the stage's own run for a single-stage report, the task's run for a whole-task one — and must match the recorded implementation commit. A newer unfinished, broken, or blocked execution prevents fallback to an older success. Prepare and `okstra handoff pr-plan` re-check this evidence. A new implementation start or completion invalidates the earlier approval.
|
|
91
91
|
- The working tree is clean.
|
|
92
92
|
- The current branch is recorded as evidence, not used as a PR head: every head comes from `okstra handoff pr-plan`. A release base branch such as `main`, `master`, `prod`, `preprod`, `staging`, or `dev` is never pushed.
|
|
93
93
|
- Each stage's `<base_commit>..<head_commit>` range is non-empty.
|
package/package.json
CHANGED
package/runtime/BUILD.json
CHANGED
|
@@ -324,8 +324,8 @@ The table below documents those prep-time seed values **for reference only** —
|
|
|
324
324
|
| Role | Seed model | Worker assignment | Source definition |
|
|
325
325
|
|------|-----------|---------------|-------------------|
|
|
326
326
|
| Lead role | opus | -- | runtime-specific role label; orchestration + convergence supervision + final-report review/approval |
|
|
327
|
-
| Report writer worker | sonnet | report-writer-worker | `
|
|
328
|
-
| Claude worker | opus | claude-worker | `
|
|
327
|
+
| Report writer worker | sonnet | report-writer-worker | common + role + duty contracts composed per invocation; `native-session` execution |
|
|
328
|
+
| Claude worker | opus | claude-worker | common + role + duty contracts composed per invocation; `native-session` execution |
|
|
329
329
|
| Codex worker | gpt-6-sol | codex-worker | duty + task instructions composed per invocation; deterministic `worker-dispatch` execution |
|
|
330
330
|
| Antigravity worker | gemini-3.1-pro | antigravity-worker | duty + task instructions composed per invocation; deterministic `worker-dispatch` execution |
|
|
331
331
|
|
|
@@ -23,7 +23,7 @@ This is where the preflight's conventions get enforced against the code you actu
|
|
|
23
23
|
|
|
24
24
|
Do not scan holistically and stop when it "looks fine". Work the matrix exhaustively — the failure mode this gate exists to prevent is a real defect surviving because you eyeballed the diff instead of enumerating it.
|
|
25
25
|
|
|
26
|
-
**Enforced:** `scripts/okstra_ctl/initial_prompt_materialization.py` `_required_resource_path_candidates` names this body for the `implementation-executor` audience, so it is appended to every executor prompt rather than left to a file reference the CLI executor cannot open
|
|
26
|
+
**Enforced:** `scripts/okstra_ctl/initial_prompt_materialization.py` `_required_resource_path_candidates` names this body for the `implementation-executor` audience, so it is appended to every executor prompt rather than left to a file reference the CLI executor cannot open. The audience selects the body, not the runner, so an in-process `native-session` worker receives the same text through the same path.
|
|
27
27
|
|
|
28
28
|
## Method — enumerate the diff, then walk every (file × applicable-rule) cell
|
|
29
29
|
|
|
@@ -27,4 +27,4 @@ fails, fix it or surface the violation — do not claim done on a failing item.
|
|
|
27
27
|
Close with a `Self-check coverage:` line naming the files you verified the
|
|
28
28
|
per-file items against, so the Coverage footer above and this gate reconcile.
|
|
29
29
|
|
|
30
|
-
**Enforced:** `scripts/okstra_ctl/initial_prompt_materialization.py` `_required_resource_path_candidates` names this body for the `implementation-executor` audience, so it is appended to every executor prompt rather than left to a file reference the CLI executor cannot open
|
|
30
|
+
**Enforced:** `scripts/okstra_ctl/initial_prompt_materialization.py` `_required_resource_path_candidates` names this body for the `implementation-executor` audience, so it is appended to every executor prompt rather than left to a file reference the CLI executor cannot open. The audience selects the body, not the runner, so an in-process `native-session` worker receives the same text through the same path.
|
|
@@ -66,7 +66,7 @@
|
|
|
66
66
|
- **Read-only command log**: any pre-existing test/validation command touched during this run MUST be listed with its exact command line and one honest status — `executed` (ran; carries its exit code) / `advisory` (external Tier 3 did not PASS; carries observed/expected results and remains user-owned) / `env-unavailable` (should run but cannot in this environment — missing replica DB, container, or service; carries the reason, never a faked pass) / `not-configured` (no such qa-command tier) / `rejected` (a mutating/denied token — skipped, carries the denied token). A check that could not run locally is recorded as `env-unavailable` or `advisory` according to the external QA policy — never silently dropped and never reported as `executed` with an invented exit code. Mutating-command prohibition is the shared read-only boundary (see Non-goals); it is not restated per row.
|
|
67
67
|
- **Could-not-verify roll-up (§5.8.9)**: the template mechanically aggregates every not-confirmed check into one scannable list — `gap` requirement-coverage rows, `advisory` / `not-configured` / `env-unavailable` / `rejected` command rows, and `blocked` manual tests. You do not hand-author it, but you MUST give those rows their honest status so nothing unverified hides across sections: a check silently recorded as `executed`/`covered` will not surface in the roll-up. This is okstra's answer to "say what could not be verified this run."
|
|
68
68
|
- **Routing recommendation**: `finalVerification.routingRecommendation` is an **object** with exactly two fields — `target`, one value of the enum below, and `rationale`, the sentence tying that choice to the verdict and the blocker list. Free routing prose is not the field; a target named only in the prose does not route the task, because Phase 7 projects `workflow.nextRecommendedPhase` from `target` alone. The eight allowed targets are `release-handoff`, `release-handoff(stage-group)`, `final-verification`, `error-analysis`, `implementation-option-selection`, `implementation-planning`, `implementation`, and `done`. `final-verification` re-runs this phase on the same head and is for exactly one situation: every remaining blocker is an environment or configuration fault whose cause this report already names — a `qaCommands` entry pointing at a path that no longer exists, a missing credential, a stale fixture — so nothing in the code, the plan, or the selected direction is being re-decided. Name the repair in the `rationale`. When any blocker needs a code, plan, or direction change, route to the phase that owns that change instead; routing a defect you have not diagnosed back into this phase re-runs the verification that already failed. Both `release-handoff` forms are allowed ONLY when the verdict is release-ready — `accepted`, or `conditional-accept` with every condition declaring `blocksReleaseHandoff: false`. Either verification scope may route there: release-handoff opens one PR per stage, so a release-ready `single-stage` run is the evidence for that stage's PR, and `release-handoff(stage-group)` is only a scope qualifier that projects onto the same phase. `done` ends the lifecycle here. Enforcement: `schemas/final-report-v2.0.schema.json` rejects a `target` outside the enum, a missing `rationale`, and a string in place of the object; `validators/validate-run.py` rejects a missing `target` and a verdict that is not release-ready routed to either `release-handoff` form (naming the condition ids that block it).
|
|
69
|
-
- **Verified-row recording** (
|
|
69
|
+
- **Verified-row recording** (both scopes): when the verdict is release-ready, the lead MUST run `okstra handoff record-verified --plan-run-root <plan-run-root> --stage <N> --report-path <final-report data.json path> --data-json <final-report data.json path>` and quote the command + exit code in the report. Pass the record path to both: the Markdown reading copy is rendered on request and does not exist in a finished run (ADR-0014), and the helper normalizes either path to the record anyway. A `whole-task` report clears every stage in its own `stageReports`, so run it **once per those stages** — each run writes that stage's row from this one report. Without those rows the stage is never offered a pull request, and release-handoff opens one PR per stage. The helper checks the latest verification manifest, task/stage identity, report pointer, prepared target, and recorded implementation commit. It records the captured commit and original verdict, including conditional acceptance conditions. A missing target or mismatched commit requires re-verification. Recording happens before final validation; eligibility is granted only after that verification passes validation. **Enforced:** `okstra_ctl.handoff_verification` validates the evidence, and `validators/validate-run.py` `_validate_verified_row_recorded` requires a `verified` row matching this report, captured commit, and verdict.
|
|
70
70
|
- Clarification request policy (phase-specific addendum — shared policy is in `_common-contract.md`):
|
|
71
71
|
- populate `## 1. Clarification Items` only when a blocker hinges on information only the user can supply (deployment intent, intended target environment, business-rule interpretation); use `Blocks=next-phase` for items that gate continuing to release-handoff
|
|
72
72
|
- Self-review pass before finalising the report (the Okstra lead runs this; do not delegate it):
|
|
@@ -7,7 +7,7 @@
|
|
|
7
7
|
- **Execution model: single-lead, no worker dispatch.** This phase is a thin orchestrator over `git` / `gh`; it does NOT dispatch teammates, does NOT dispatch analysis or drafter sub-agents, and does NOT run convergence. The host-native Okstra lead performs every step inline (drafting PR text, asking the user, running git / gh, writing the final report) — see "Lead-only contract" below.
|
|
8
8
|
- Worker roster: none — this profile intentionally has no `- Required workers:` block; the run is executed entirely by the Okstra lead.
|
|
9
9
|
- Lead-only contract (replaces the shared team contract for this phase):
|
|
10
|
-
- The host-native Okstra lead is the sole agent for this run. No worker dispatch, no teammates, no parallel sub-agents, no convergence loop.
|
|
10
|
+
- The host-native Okstra lead is the sole agent for this run. No worker dispatch, no teammates, no parallel sub-agents, no convergence loop. Do NOT run `okstra convergence` in any form: prepare already wrote this run's convergence state as "not run — fewer than two analysers", which is the record report assembly reads. **Enforced:** `okstra_ctl.render._initialize_lead_only_convergence` writes it when the run has no analysers, and leaves an existing state alone.
|
|
11
11
|
- The lead drafts each stage's PR title and PR body **inline** by reading the run brief, that stage's cited final-verification report, `git log --oneline <stage base>..<stage head>`, and `git diff <stage base>..<stage head> --stat`. No drafter worker is dispatched.
|
|
12
12
|
- The lead authors the final-report file directly (no `Report writer worker` dispatch). The report still conforms to the standard `templates/reports/final-report-v2.template.md` structure, including the `## 5.6 Release Handoff Deliverables` section.
|
|
13
13
|
- The shared anti-escalation rule from the common contract still applies: do not start any other lifecycle phase from inside this run.
|
|
@@ -20,7 +20,7 @@
|
|
|
20
20
|
- the lead MUST capture `git status --short` and confirm the working tree is clean. Dirty state aborts the run; release-handoff packages the commits produced by `implementation`, it does not stage or commit changes.
|
|
21
21
|
- the lead MUST capture `git rev-parse --abbrev-ref HEAD` and record it as the run's current branch. It is evidence, not a PR head: every PR head comes from `okstra handoff pr-plan`. If the current branch is `main`, `master`, `prod`, `preprod`, `staging`, or `dev`, record it and do not push it — release-handoff never pushes a base branch.
|
|
22
22
|
- before pushing, run `okstra handoff pr-plan` again if anything mutated since entry, and confirm each selected stage is still present in its output. `pr-plan` re-checks eligibility and each stage's recorded verified commit. Record the pr-plan output in Executed Commands. These pre-push comparisons are lead instructions; the automated entry checks are enforced by `okstra_ctl.handoff_verification` and the handoff tests.
|
|
23
|
-
- **Commit history is preserved — structurally, not as a preference.** These PRs are a stack: stage N's PR base is stage N-1's branch. Squash-merging a PR in the stack replaces its commits with a new one, so the next PR's base points at commits that no longer exist on the target branch and its diff re-shows the predecessor's changes. Therefore: okstra never rebases, squashes, amends, or cherry-picks a stage branch, and
|
|
23
|
+
- **Commit history is preserved — structurally, not as a preference.** These PRs are a stack: stage N's PR base is stage N-1's branch. Squash-merging a PR in the stack replaces its commits with a new one, so the next PR's base points at commits that no longer exist on the target branch and its diff re-shows the predecessor's changes. Therefore: okstra never rebases, squashes, amends, or cherry-picks a stage branch, and the merge policy — **merge commit or rebase-merge, never squash, and merge in ascending stage order** — is stated to the operator in the final report's `Stage PR Plan` and next-action narrative. It does NOT go into the PR body: the reviewer of one stage's PR is not the person merging the stack, and the block was noise they deleted by hand (2026-09-24, dev-10860 PRs #1487-#1490).
|
|
24
24
|
- User interaction protocol (Okstra lead — performed in order, using the selected runtime adapter's interactive prompt):
|
|
25
25
|
1. **Action selection** — present three choices and capture exactly one:
|
|
26
26
|
- `local checkout` — bring one stage's branch into the MAIN worktree for local testing (no push, no PR). See step 1c.
|
|
@@ -39,6 +39,7 @@
|
|
|
39
39
|
`okstra handoff pr-plan --plan-run-root <...> --approved-plan <...> --project-root <project root> --project-id <id> --task-group <g> --task-id <t> --work-category <c> --stages <HANDOFF_STAGES> --base <chosen-base>`.
|
|
40
40
|
It returns one row per stage with `head_branch`, `head_commit`, `base_kind` (`release-base` | `stage` | `merge-base`), `base_branch` and `base_commit`. Exit 2 means the predecessors of a multi-dependency stage conflict with each other: show the `conflicts` paths and stop (route: resolve manually, then retry). Exit 1 means an eligibility or dependency violation — most often a stage whose predecessor is neither in this run nor already PR'd nor merged into the release base: show the error verbatim and re-ask the action selection. On success these rows are authoritative for every subsequent step; the lead does not recompute a base by hand.
|
|
41
41
|
A `merge-base` row means `pr-plan` created a branch merging that stage's predecessors. That branch is a PR base only — it is never a PR head — and it MUST be pushed before the PR that targets it.
|
|
42
|
+
The output also carries `push_plan` and `warnings`. `push_plan` is every branch the PRs sit on, already in push order (each `merge-base` first, then the stage heads in ascending order), with the `commit` it must reach and an `action` of `create` (origin does not have it), `fast-forward` (origin is behind) or `up-to-date` (origin already has that commit — do not push it). GitHub resolves a PR's head and base by name on origin, so this list, not the local branch set, is what makes the stack real. A `warnings` row means `pr-plan` fast-forwarded a local branch to origin because origin was ahead: the commits origin adds are outside what this run verified. Quote every `warnings` row verbatim in Executed Commands and say so in the branch/PR explanation — the user is deciding whether to release commits this run did not verify. `pr-plan` refuses outright when a branch moved outside okstra or diverged from origin, since only a force push could deliver it.
|
|
42
43
|
2b. **Pre-merge conflict probe** (only when the user picked `push + PR`) — before pushing, for each planned stage in ascending order:
|
|
43
44
|
- run `git fetch origin <chosen-base>` once (read-only on the local working tree).
|
|
44
45
|
- run `git merge-tree --write-tree <head_branch> <base_branch>` for that stage, using the `base_branch` from the pr-plan row (`origin/<chosen-base>` for a `release-base` row). Exit 0 means clean. Exit 1 with a result-tree object ID on the first stdout line means a merge conflict. A missing result tree, or any other exit code, is an execution error: stop and report it, without offering permission to proceed as if it were a content conflict. An invalid ref can also return exit 1, so the exit code alone is insufficient.
|
|
@@ -57,12 +58,12 @@
|
|
|
57
58
|
- **PR body template** — the run context exposes `PR_TEMPLATE_PATH` and `PR_TEMPLATE_SOURCE`. The path MUST be an okstra-owned project artifact under `<PROJECT_ROOT>/.okstra/**`, or a file the prepare step already materialised into this run's artifact directory. If the resolved file is missing or outside that boundary at draft time, abort with a clear error — do NOT invent a structure. Otherwise the lead MUST `Read` it verbatim, strip HTML comments, and fill in the placeholders, treating the template (never a hard-coded section list) as the source of truth for the structure.
|
|
58
59
|
- produce **two artifacts per stage** before showing them to the user:
|
|
59
60
|
1. **PR title** — by default the subject of that stage's most recent implementation commit, or a concise Conventional Commits-style summary of the stage's committed range.
|
|
60
|
-
2. **PR body** — markdown filled from `PR_TEMPLATE_PATH`,
|
|
61
|
+
2. **PR body** — markdown filled from `PR_TEMPLATE_PATH`, and nothing else. Do NOT append a stack plan, a stage checklist, a sentence about where this PR stops in the stack, or a merge-policy line: the stack's shape and merge rule belong to the final report and the closeout, not to a reviewer's PR. The user-confirmation step's diff (Q3 `edit then proceed`) is computed against the filled template, not against the raw template file.
|
|
61
62
|
- **Filled-body self-check (runs on each drafted body, before Q3 shows it).** "Fill in the placeholders" is only observable if someone checks that they were filled, so scan the drafted body for these and carry the result into Q3: a residual HTML comment or template marker (`<!-- … -->`, `_Describe your changes…_`, a bare `TODO` / `FIXME` the diff did not introduce), an unchecked `- [ ]` box with no inline N/A justification, a section heading whose body is empty, and a whole body short enough to carry no substance. Each hit is reported to the user with its line — never silently left in, and never auto-filled with invented content. The user remains free to accept it; the point is that they see it before choosing `use as-is`.
|
|
62
63
|
- Allowed actions during the run (Okstra lead only):
|
|
63
64
|
- read-only inspection: `git status`, `git status --short`, `git diff`, `git log`, `git rev-parse`, `git ls-remote --heads origin <name>`, `gh pr list --head <branch>`, `gh pr view <url>`.
|
|
64
65
|
- merge-conflict probe (only when the user picked `push + PR`): `git fetch origin <chosen-base>` and the exact `git merge-tree` commands from step 2b. Both preserve the working tree and branch tips.
|
|
65
|
-
- branch push (only when the user picked `push + PR`)
|
|
66
|
+
- branch push (only when the user picked `push + PR`): `git push -u origin <branch>` for each `push_plan` row whose `action` is `create` or `fast-forward`, in the order `push_plan` lists them (each `merge-base` branch before the PR that targets it, then the stage heads in ascending order). Skip an `up-to-date` row — origin already holds that commit. The pushed ref MUST come from a `push_plan` row — never the release base branch, and never a ref that list does not name.
|
|
66
67
|
- PR creation, one per stage (only when the user picked `push + PR` AND no PR with the same head already exists on origin): `gh pr create --base <row base_branch> --head <row head_branch> --title "<title>" --body "<body>"`. The title and body are the user-confirmed PR draft for that stage. Open them in ascending stage order so each base branch already exists on origin.
|
|
67
68
|
- PR reuse: if `gh pr list --head <branch> --state open --json url --jq '.[0].url'` returns a URL, treat that PR as already existing — record the URL in the final report and SKIP `gh pr create` for that stage.
|
|
68
69
|
- after each `gh pr create` succeeds (or an existing PR is reused), the lead MUST run `okstra handoff record-pr --plan-run-root <...> --stage <N> --branch <head branch> --base <base branch> --url <pr url>` and quote the command + exit code in the final report. One record-pr call per stage — the consumers ledger is what stops the same stage being PR'd twice.
|
|
@@ -72,7 +73,8 @@
|
|
|
72
73
|
- Required deliverable shape (final report, in addition to the standard sections):
|
|
73
74
|
- **Source Verification Reports**: one row per selected stage — relative path of that stage's originating `final-verification` final-report plus the literal quoted `Verdict Token` row, and — when that token is `conditional-accept` — every Conditional Acceptance Condition row quoted with its `blocksReleaseHandoff` value.
|
|
74
75
|
- **Feature Branch & Working-Tree State**: the branch from `git rev-parse --abbrev-ref HEAD` at run start and the output of `git status --short` at run start.
|
|
75
|
-
- **Stage PR Plan**: the pr-plan rows — stage, head branch, head commit, base kind, base branch, base commit — plus, for every `merge-base` row, which predecessor stages it merges and the merge commits it created. This is the table a reader uses to merge the stack in the right order.
|
|
76
|
+
- **Stage PR Plan**: the pr-plan rows — stage, head branch, head commit, base kind, base branch, base commit — plus, for every `merge-base` row, which predecessor stages it merges and the merge commits it created. This is the table a reader uses to merge the stack in the right order, so it carries the merge rule with it: merge in ascending stage order, with a merge commit or rebase-merge, never a squash.
|
|
77
|
+
- **User narrative `nextAction`**: when this run opened or reused any PR, it names the merge order (ascending stage number) and the merge rule (merge commit or rebase-merge, never squash), and says why — squashing one PR in the stack invalidates the next PR's base. This is where the operator reads the rule; the PR bodies do not carry it.
|
|
76
78
|
- **User Selections**: a block recording each prompt and the user's verbatim answer.
|
|
77
79
|
- Q1 action: `local checkout` | `push + PR` | `skip`.
|
|
78
80
|
- Q1c checkout stage (only when Q1 is `local checkout`): the `--stage <N>` value this answer produced. Record it in the `User Selections` row `H1c` (`userSelections.h1c`); omit the row entirely when Q1 was not `local checkout`.
|
|
@@ -101,7 +103,7 @@
|
|
|
101
103
|
2. **User-selection traceability** — every executed mutating command maps to a user selection captured in the report. Any mutating command without a corresponding user answer is a contract violation.
|
|
102
104
|
3. **Forbidden-action audit** — scan the run's session transcripts (`git`, `gh` invocations) for every entry in the Forbidden actions list above. Any occurrence means the run has crossed into unsafe territory and MUST be flagged as `contract-violated`.
|
|
103
105
|
4. **Push-target audit** — for every `git push` recorded, confirm the refspec resolves to a branch that appears in the pr-plan output, not to the release base branch.
|
|
104
|
-
5. **Stack-integrity audit** — for every PR opened, confirm its base is the `base_branch` of that stage's pr-plan row
|
|
106
|
+
5. **Stack-integrity audit** — for every PR opened, confirm its base is the `base_branch` of that stage's pr-plan row and that the base branch was pushed first. A stacked PR opened against the release base instead of its predecessor ships the predecessor's diff twice. Confirm too that no PR body carries a stack plan or merge-policy block — that text belongs to the report and the closeout.
|
|
105
107
|
6. **Idempotency check** — for every stage whose PR already existed at run start, confirm the report records `PR reused` rather than a fresh `gh pr create` invocation.
|
|
106
108
|
7. **Merge-conflict probe audit** — for any `push + PR` run, confirm the report's `Merge Conflict Probe` section is present and either records `Clean` or records `Conflicts detected` with the user's verbatim choice. A missing or unparseable probe entry on a `push + PR` run is a contract violation.
|
|
107
109
|
- Non-goals:
|
|
@@ -193,7 +193,8 @@
|
|
|
193
193
|
"echo_template": "handoff-stages: {value}",
|
|
194
194
|
"labels": {
|
|
195
195
|
"stage": "stage {stage} (depends-on: {deps})",
|
|
196
|
-
"blocked_none": "없음"
|
|
196
|
+
"blocked_none": "없음",
|
|
197
|
+
"all_stages": "전체 stage ({stages}) — 자격 있는 stage 마다 PR 하나"
|
|
197
198
|
},
|
|
198
199
|
"echo_variants": {
|
|
199
200
|
"stages": "handoff stages: {stages}"
|
|
@@ -201,7 +202,7 @@
|
|
|
201
202
|
"errors": {
|
|
202
203
|
"no_plan": "release-handoff 는 이 task 의 implementation-planning final-report 가 필요합니다 — implementation-planning → implementation → final-verification 을 먼저 진행하세요.",
|
|
203
204
|
"plan_not_approved": "최신 plan 이 approved 상태가 아닙니다: {plan} — implementation 이 완료된 task 에서만 release-handoff 를 시작할 수 있습니다.",
|
|
204
|
-
"nothing_eligible": "PR 로 내보낼 수 있는 stage 가 없습니다 — 제외 사유: {blocked}.
|
|
205
|
+
"nothing_eligible": "PR 로 내보낼 수 있는 stage 가 없습니다 — 제외 사유: {blocked}. final-verification 의 accepted 판정을 `okstra handoff record-verified` 로 먼저 기록하세요 — 단독-stage 판정과 전체 task 판정 모두 근거가 됩니다.",
|
|
205
206
|
"none_selected": "최소 1개의 stage 를 선택하세요",
|
|
206
207
|
"not_eligible": "eligible 하지 않은 stage: {bad} (선택 가능: {eligible})"
|
|
207
208
|
}
|
|
@@ -218,9 +218,15 @@ def _append_row(plan_run_root: Path, record: Dict[str, Any]) -> None:
|
|
|
218
218
|
|
|
219
219
|
def append_verified(plan_run_root: Path, *, impl_task_key: str, stage: int,
|
|
220
220
|
verdict: str, report_path: str, head_commit: str = "",
|
|
221
|
-
final_verdict: dict | None = None
|
|
222
|
-
|
|
223
|
-
|
|
221
|
+
final_verdict: dict | None = None,
|
|
222
|
+
verification_scope: str = "",
|
|
223
|
+
verified_head_commit: str = "") -> None:
|
|
224
|
+
"""final-verification 결과를 stage 별로 기록. 같은 report_path 재기록은 멱등,
|
|
225
|
+
다른 report_path 는 재검증으로 append 한다 (읽기는 last-wins).
|
|
226
|
+
|
|
227
|
+
`head_commit` 은 언제나 이 stage 가 완료로 기록한 커밋이다. 전체 task 검증이
|
|
228
|
+
근거인 행은 검증된 head 가 그 커밋이 아니므로, 그 head 를
|
|
229
|
+
`verified_head_commit` 으로 함께 남긴다 — 자격 판정이 둘을 각각 대조한다."""
|
|
224
230
|
with consumers_mutex(plan_run_root):
|
|
225
231
|
rows = read_consumers(plan_run_root)
|
|
226
232
|
done = latest_done_by_stage(rows).get(stage, {})
|
|
@@ -234,13 +240,19 @@ def append_verified(plan_run_root: Path, *, impl_task_key: str, stage: int,
|
|
|
234
240
|
and row.get("status") == "verified"
|
|
235
241
|
and row.get("report_path") == report_path):
|
|
236
242
|
return
|
|
237
|
-
|
|
243
|
+
row: Dict[str, Any] = {
|
|
238
244
|
"impl_task_key": impl_task_key, "stage": stage,
|
|
239
245
|
"status": "verified", "verdict": verdict,
|
|
240
246
|
"report_path": report_path,
|
|
241
247
|
"head_commit": head_commit,
|
|
242
248
|
"final_verdict": final_verdict or {},
|
|
243
|
-
}
|
|
249
|
+
}
|
|
250
|
+
# 단독-stage 행에는 싣지 않는다 — 그 행의 검증된 head 는 `head_commit`
|
|
251
|
+
# 자신이고, 빈 필드를 더하면 옛 행과 새 행의 모양이 갈라진다.
|
|
252
|
+
if verification_scope:
|
|
253
|
+
row["verification_scope"] = verification_scope
|
|
254
|
+
row["verified_head_commit"] = verified_head_commit
|
|
255
|
+
_append_row(plan_run_root, row)
|
|
244
256
|
|
|
245
257
|
|
|
246
258
|
def append_pr(plan_run_root: Path, *, impl_task_key: str, stages: List[int],
|
|
@@ -849,6 +849,44 @@ def validate_working_state(state: Mapping[str, Any]) -> list[str]:
|
|
|
849
849
|
return errors
|
|
850
850
|
|
|
851
851
|
|
|
852
|
+
def lead_only_final_state(task_key: str) -> dict[str, Any]:
|
|
853
|
+
"""워커를 띄우지 않는 phase 의 수렴 최종 상태.
|
|
854
|
+
|
|
855
|
+
리포트 계약 3.0 은 모든 리포트에 수렴 입력을 요구한다(ADR-0020: 사실마다
|
|
856
|
+
소유자가 하나). release-handoff 처럼 프로필이 "convergence 를 돌지 않는다"
|
|
857
|
+
고 선언한 phase 에는 그 파일을 만드는 주체가 없었고, report-finalize 가
|
|
858
|
+
`required input is missing` 으로 멈췄다. 리드는 매 run 마다 seed →
|
|
859
|
+
plan-round → finalize 를 손으로 돌려 빈 상태를 만들어 왔다(실측: dev-10860
|
|
860
|
+
release-handoff 001·002 둘 다).
|
|
861
|
+
|
|
862
|
+
"수렴이 돌지 않았다" 는 부재가 아니라 사실이므로 파일로 남긴다 — 분석자가
|
|
863
|
+
둘 미만이어서 자동 비활성이라는 뜻이고, 그게 이 어휘의 `auto-disabled` 다.
|
|
864
|
+
"""
|
|
865
|
+
state = {
|
|
866
|
+
"schemaVersion": FINAL_SCHEMA_VERSION,
|
|
867
|
+
"taskKey": task_key,
|
|
868
|
+
"config": {
|
|
869
|
+
"enabled": False,
|
|
870
|
+
"adversarial": False,
|
|
871
|
+
"maxRounds": 1,
|
|
872
|
+
"effectiveMaxRounds": 1,
|
|
873
|
+
"verificationMode": "lightweight",
|
|
874
|
+
"autoDisabled": "fewer-than-two-analysers",
|
|
875
|
+
},
|
|
876
|
+
"findings": [],
|
|
877
|
+
"roundHistory": [],
|
|
878
|
+
"round2SkippedReason": "auto-disabled",
|
|
879
|
+
"finalState": "converged",
|
|
880
|
+
"totalRounds": 0,
|
|
881
|
+
"finalClassificationCounts": classification_counts_from_findings([]),
|
|
882
|
+
}
|
|
883
|
+
errors = validate_final_state(state)
|
|
884
|
+
if errors:
|
|
885
|
+
raise ConvergenceContractError(
|
|
886
|
+
"lead-only convergence state is invalid: " + "; ".join(errors))
|
|
887
|
+
return state
|
|
888
|
+
|
|
889
|
+
|
|
852
890
|
def validate_final_state(state: Mapping[str, Any]) -> list[str]:
|
|
853
891
|
"""Validate historical v1.0-v1.2 and strict current v1.3 artifacts."""
|
|
854
892
|
errors: list[str] = []
|
|
@@ -33,6 +33,10 @@ from .worktree import (compute_branch_name, main_worktree_path,
|
|
|
33
33
|
is_ancestor, remove_worktree_force, _git)
|
|
34
34
|
|
|
35
35
|
|
|
36
|
+
SINGLE_STAGE_SCOPE = "single-stage"
|
|
37
|
+
WHOLE_TASK_SCOPE = "whole-task"
|
|
38
|
+
|
|
39
|
+
|
|
36
40
|
class HandoffError(Exception):
|
|
37
41
|
"""자격/전제 위반 — exit 1, actionable 메시지."""
|
|
38
42
|
|
|
@@ -58,6 +62,9 @@ def compute_eligibility(stage_map: List[Dict[str, Any]],
|
|
|
58
62
|
stage_map,
|
|
59
63
|
rows,
|
|
60
64
|
).handoff_eligibility()
|
|
65
|
+
# 재구현된 stage 의 커밋 동일성은 `consumers.verified_accepted_stages` 가
|
|
66
|
+
# 이미 본다(`handoff_eligibility` 가 그 집합을 읽는다). 여기 남는 판정은
|
|
67
|
+
# "그 승인 행이 가리키는 검증 run 이 아직 최신인가" 하나다.
|
|
61
68
|
last = {row["stage"]: row for row in rows if row.get("status") == "verified"}
|
|
62
69
|
for item in eligibility:
|
|
63
70
|
if item["eligible"] and not verification_row_is_current(last.get(item["stage"], {})):
|
|
@@ -175,6 +182,147 @@ def _resolve_stage_base(
|
|
|
175
182
|
"base_branch_created": merged.created}
|
|
176
183
|
|
|
177
184
|
|
|
185
|
+
def _local_tip(project_root, branch: str) -> str:
|
|
186
|
+
"""로컬 브랜치의 현재 커밋. 없으면 빈 문자열."""
|
|
187
|
+
return _git(Path(project_root), "rev-parse", "--verify", "--quiet",
|
|
188
|
+
f"refs/heads/{branch}").stdout.strip()
|
|
189
|
+
|
|
190
|
+
|
|
191
|
+
def _remote_tips(project_root, branches: List[str]) -> Dict[str, str]:
|
|
192
|
+
"""origin 이 지금 가진 브랜치와 커밋. 없는 브랜치는 키가 없다.
|
|
193
|
+
|
|
194
|
+
`ls-remote` 로 묻는다 — 비교하려고 ref 를 쓰지 않고, 없는 브랜치를 fetch
|
|
195
|
+
하려다 실패하지도 않는다.
|
|
196
|
+
"""
|
|
197
|
+
if not branches:
|
|
198
|
+
return {}
|
|
199
|
+
res = _git(Path(project_root), "ls-remote", "--heads", "origin", *branches)
|
|
200
|
+
if res.returncode != 0:
|
|
201
|
+
raise HandoffError(f"cannot read origin branch state: {res.stderr.strip()}")
|
|
202
|
+
tips: Dict[str, str] = {}
|
|
203
|
+
for line in res.stdout.splitlines():
|
|
204
|
+
parts = line.split()
|
|
205
|
+
if len(parts) == 2 and parts[1].startswith("refs/heads/"):
|
|
206
|
+
tips[parts[1][len("refs/heads/"):]] = parts[0]
|
|
207
|
+
return tips
|
|
208
|
+
|
|
209
|
+
|
|
210
|
+
def _branch_worktree(project_root, branch: str) -> Optional[Path]:
|
|
211
|
+
"""그 브랜치를 체크아웃해 둔 워크트리. 없으면 None.
|
|
212
|
+
|
|
213
|
+
체크아웃된 브랜치는 `git branch -f` 로 옮길 수 없으므로, fast-forward 는 그
|
|
214
|
+
워크트리 안에서 해야 한다.
|
|
215
|
+
"""
|
|
216
|
+
res = _git(Path(project_root), "worktree", "list", "--porcelain")
|
|
217
|
+
if res.returncode != 0:
|
|
218
|
+
return None
|
|
219
|
+
current: Optional[Path] = None
|
|
220
|
+
for line in res.stdout.splitlines():
|
|
221
|
+
if line.startswith("worktree "):
|
|
222
|
+
current = Path(line[len("worktree "):].strip())
|
|
223
|
+
elif line.strip() == f"branch refs/heads/{branch}":
|
|
224
|
+
return current
|
|
225
|
+
return None
|
|
226
|
+
|
|
227
|
+
|
|
228
|
+
def _fast_forward_local(project_root, branch: str, remote_commit: str) -> None:
|
|
229
|
+
"""로컬 브랜치를 origin 의 커밋으로 fast-forward 한다.
|
|
230
|
+
|
|
231
|
+
호출자가 조상 관계를 이미 확인했으므로 내용이 사라질 수 없다. 체크아웃
|
|
232
|
+
상태면 그 워크트리에서 `merge --ff-only` 로, 아니면 ref 를 옮긴다.
|
|
233
|
+
"""
|
|
234
|
+
_run_git(["fetch", "origin", branch], project_root)
|
|
235
|
+
holder = _branch_worktree(project_root, branch)
|
|
236
|
+
if holder is None:
|
|
237
|
+
_run_git(["branch", "-f", branch, remote_commit], project_root)
|
|
238
|
+
return
|
|
239
|
+
if is_dirty_excluding_okstra(holder):
|
|
240
|
+
raise HandoffError(
|
|
241
|
+
f"branch {branch} is behind origin and is checked out in {holder}, "
|
|
242
|
+
"which has uncommitted changes — commit or stash them so the "
|
|
243
|
+
"branch can be fast-forwarded to origin")
|
|
244
|
+
_run_git(["merge", "--ff-only", remote_commit], holder)
|
|
245
|
+
|
|
246
|
+
|
|
247
|
+
# 계획 행이 부르는 브랜치를 실제로 푸시할 순서. merge-base 는 그것을 base 로
|
|
248
|
+
# 삼는 PR 보다 먼저 origin 에 있어야 한다(프로필 §"branch push" 와 같은 순서).
|
|
249
|
+
def _plan_branch_order(planned: List[Dict[str, Any]]) -> List[tuple[str, str]]:
|
|
250
|
+
ordered: List[tuple[str, str]] = []
|
|
251
|
+
seen: set[str] = set()
|
|
252
|
+
for row in planned:
|
|
253
|
+
if row["base_kind"] == "merge-base" and row["base_branch"] not in seen:
|
|
254
|
+
seen.add(row["base_branch"])
|
|
255
|
+
ordered.append((row["base_branch"], row["base_commit"]))
|
|
256
|
+
for row in planned:
|
|
257
|
+
if row["head_branch"] not in seen:
|
|
258
|
+
seen.add(row["head_branch"])
|
|
259
|
+
ordered.append((row["head_branch"], row["head_commit"]))
|
|
260
|
+
return ordered
|
|
261
|
+
|
|
262
|
+
|
|
263
|
+
def _reconcile_with_origin(
|
|
264
|
+
project_root, planned: List[Dict[str, Any]],
|
|
265
|
+
) -> tuple[List[Dict[str, Any]], List[str]]:
|
|
266
|
+
"""PR 이 올라앉을 브랜치를 origin 기준으로 맞춘다.
|
|
267
|
+
|
|
268
|
+
PR 의 head 와 base 는 GitHub 이 **이름으로 origin 에서** 찾는다. 계획이
|
|
269
|
+
로컬 ref 만 보고 짜이던 동안, origin 의 같은 이름이 다른 커밋에 있어도
|
|
270
|
+
계획은 그 사실을 몰랐다 — push 단계에서 거절되거나, PR 의 diff 가 계획서의
|
|
271
|
+
커밋과 달랐다.
|
|
272
|
+
|
|
273
|
+
네 가지 상태를 가른다. origin 에 없거나 뒤처져 있으면 push 가 해결하므로
|
|
274
|
+
무엇을 어떤 순서로 올려야 하는지만 계획에 싣는다. origin 이 앞서 있으면
|
|
275
|
+
로컬을 fast-forward 하고, 그 추가 커밋이 이 run 이 검증한 범위 밖이라는
|
|
276
|
+
사실을 경고로 남긴다. 갈라져 있으면 거절한다 — 그 상태를 push 로 넘기는
|
|
277
|
+
길은 force 뿐이고, 그건 열린 PR 의 리뷰 커밋을 지운다.
|
|
278
|
+
"""
|
|
279
|
+
order = _plan_branch_order(planned)
|
|
280
|
+
tips = _remote_tips(project_root, [branch for branch, _ in order])
|
|
281
|
+
push_plan: List[Dict[str, Any]] = []
|
|
282
|
+
warnings: List[str] = []
|
|
283
|
+
for branch, recorded in order:
|
|
284
|
+
local = _local_tip(project_root, branch)
|
|
285
|
+
if not local:
|
|
286
|
+
raise HandoffError(
|
|
287
|
+
f"branch {branch} does not exist in this repository — there is "
|
|
288
|
+
"nothing to open a PR from")
|
|
289
|
+
# tip 이 기록된 커밋보다 앞선 것은 정상 상태다 — git-reconcile 이 done
|
|
290
|
+
# 이후의 커밋을 그 브랜치에 남긴다. 담기지 않은 것만 거절한다: 검증한
|
|
291
|
+
# 커밋이 브랜치에 없으면 그 PR 은 검증된 것을 나르지 않는다.
|
|
292
|
+
if recorded and local != recorded:
|
|
293
|
+
if not is_ancestor(project_root, recorded, local):
|
|
294
|
+
raise HandoffError(
|
|
295
|
+
f"branch {branch} is at {local}, which does not contain the "
|
|
296
|
+
f"commit this task's ledger verified for it ({recorded}) — "
|
|
297
|
+
"the branch was rewritten, so its PR would not carry the "
|
|
298
|
+
"verified work")
|
|
299
|
+
warnings.append(
|
|
300
|
+
f"branch {branch} is ahead of the commit this run verified "
|
|
301
|
+
f"({recorded} -> {local}); the commits it adds are outside the "
|
|
302
|
+
"verified scope")
|
|
303
|
+
remote = tips.get(branch)
|
|
304
|
+
if remote is None:
|
|
305
|
+
action = "create"
|
|
306
|
+
elif remote == local:
|
|
307
|
+
action = "up-to-date"
|
|
308
|
+
elif is_ancestor(project_root, remote, local):
|
|
309
|
+
action = "fast-forward"
|
|
310
|
+
elif is_ancestor(project_root, local, remote):
|
|
311
|
+
_fast_forward_local(project_root, branch, remote)
|
|
312
|
+
action = "up-to-date"
|
|
313
|
+
warnings.append(
|
|
314
|
+
f"branch {branch} was behind origin ({local} -> {remote}) and "
|
|
315
|
+
"was fast-forwarded to it; the commits origin adds are outside "
|
|
316
|
+
f"what this run verified ({recorded or local})")
|
|
317
|
+
else:
|
|
318
|
+
raise HandoffError(
|
|
319
|
+
f"branch {branch} and origin/{branch} have diverged (local "
|
|
320
|
+
f"{local}, origin {remote}) — only a force push could deliver "
|
|
321
|
+
"this branch, which would drop the commits origin holds")
|
|
322
|
+
push_plan.append({"branch": branch, "commit": local, "action": action})
|
|
323
|
+
return push_plan, warnings
|
|
324
|
+
|
|
325
|
+
|
|
178
326
|
def pr_plan(*, project_root, plan_run_root, stage_map, stages, base_branch,
|
|
179
327
|
work_category, project_id, task_group, task_id) -> Dict[str, Any]:
|
|
180
328
|
"""선택한 stage 각각의 PR head/base 를 정하고, 필요한 선행 머지 브랜치를
|
|
@@ -210,7 +358,9 @@ def pr_plan(*, project_root, plan_run_root, stage_map, stages, base_branch,
|
|
|
210
358
|
anchor_base=anchor_base),
|
|
211
359
|
}
|
|
212
360
|
planned.append(row)
|
|
213
|
-
|
|
361
|
+
push_plan, warnings = _reconcile_with_origin(project_root, planned)
|
|
362
|
+
return {"ok": True, "release_base": base_branch, "stages": planned,
|
|
363
|
+
"push_plan": push_plan, "warnings": warnings}
|
|
214
364
|
|
|
215
365
|
|
|
216
366
|
def _remove_task_worktree(main_wt, target: str) -> Optional[str]:
|
|
@@ -313,9 +463,18 @@ def _impl_task_key_for(rows: List[Dict[str, Any]], stage: int) -> str:
|
|
|
313
463
|
|
|
314
464
|
def record_verified(*, plan_run_root, stage: int, report_path: str,
|
|
315
465
|
data_json) -> Dict[str, Any]:
|
|
316
|
-
"""릴리스로 넘어갈 수 있는
|
|
317
|
-
|
|
318
|
-
|
|
466
|
+
"""릴리스로 넘어갈 수 있는 검증 판정을 stage 별로 기록.
|
|
467
|
+
|
|
468
|
+
범위는 둘 다 받는다. 단독-stage 판정은 그 stage 만의 근거이고, 전체 task
|
|
469
|
+
판정은 자기 `stageReports` 에 실린 stage 들의 근거다 — 검증 대상이 tip
|
|
470
|
+
stage 의 브랜치이므로(설계 결정 7) 그 브랜치는 완료된 다른 stage 의 커밋을
|
|
471
|
+
담고 있고, 그 사실은 커밋 조상 관계로 확인한다. 전체 task 로 검증한 task 가
|
|
472
|
+
stage 별 자격을 얻을 길이 없던 동안, 모든 stage 를 검증해 놓고도 stage 마다
|
|
473
|
+
단독 검증을 다시 돌려야 PR 을 열 수 있었다(2026-09-24, dev-10860).
|
|
474
|
+
|
|
475
|
+
data.json 의 taskType/scope/verdict 를 검증해 lead 가 임의 보고서를
|
|
476
|
+
verified 로 올리는 것을 막는다. 기록되는 verdict 는 보고서가 실제로 실은
|
|
477
|
+
토큰이다."""
|
|
319
478
|
try:
|
|
320
479
|
data = load_owned_object(Path(data_json), artifact="final verification report")
|
|
321
480
|
except JsonBoundaryError as exc:
|
|
@@ -323,10 +482,10 @@ def record_verified(*, plan_run_root, stage: int, report_path: str,
|
|
|
323
482
|
if (data.get("header") or {}).get("taskType") != "final-verification":
|
|
324
483
|
raise HandoffError("data.json is not a final-verification report")
|
|
325
484
|
scope = data.get("verificationScope")
|
|
326
|
-
if scope
|
|
485
|
+
if scope not in (SINGLE_STAGE_SCOPE, WHOLE_TASK_SCOPE):
|
|
327
486
|
raise HandoffError(
|
|
328
|
-
f"record-verified requires verificationScope
|
|
329
|
-
f"got {scope!r}")
|
|
487
|
+
f"record-verified requires verificationScope "
|
|
488
|
+
f"{SINGLE_STAGE_SCOPE} or {WHOLE_TASK_SCOPE}, got {scope!r}")
|
|
330
489
|
token = verdict_token(data)
|
|
331
490
|
if not release_handoff_allowed(data):
|
|
332
491
|
blocking = blocking_condition_ids(data)
|
|
@@ -339,17 +498,23 @@ def record_verified(*, plan_run_root, stage: int, report_path: str,
|
|
|
339
498
|
f"condition declares `blocksReleaseHandoff: false` — {detail}")
|
|
340
499
|
rows = consumers.read_consumers(Path(plan_run_root))
|
|
341
500
|
key = _impl_task_key_for(rows, stage)
|
|
342
|
-
_record_verified_target(Path(plan_run_root), stage, key, report_path,
|
|
501
|
+
_record_verified_target(Path(plan_run_root), stage, key, report_path,
|
|
502
|
+
Path(data_json), data, rows)
|
|
343
503
|
return {"ok": True, "stage": stage, "report_path": report_path}
|
|
344
504
|
|
|
345
505
|
|
|
346
506
|
def _record_verified_target(plan_run_root: Path, stage: int, key: str,
|
|
347
|
-
report_path: str, data_json: Path, data: dict
|
|
507
|
+
report_path: str, data_json: Path, data: dict,
|
|
508
|
+
rows: List[Dict[str, Any]]) -> None:
|
|
348
509
|
"""검증한 실행과 커밋을 입증한 뒤 같은 원장 잠금 안에서 기록한다."""
|
|
349
510
|
try:
|
|
350
511
|
ref = RunRef.from_run_dir(Path(plan_run_root).resolve()).sibling("final-verification")
|
|
351
|
-
|
|
352
|
-
|
|
512
|
+
scope = str(data.get("verificationScope") or "")
|
|
513
|
+
# 전체 task 검증은 stage 범위 run 이 아니라 task 범위 run 에 산다.
|
|
514
|
+
verification_ref = RunRef.from_task_root(
|
|
515
|
+
ref.task_root, ref.task_type,
|
|
516
|
+
stage=stage if scope == SINGLE_STAGE_SCOPE else None)
|
|
517
|
+
latest, manifest, verified_data = load_latest_verification(verification_ref.run_dir)
|
|
353
518
|
supplied = final_report_data_path(Path(data_json).resolve())
|
|
354
519
|
report_arg = Path(report_path)
|
|
355
520
|
if not report_arg.is_absolute():
|
|
@@ -357,16 +522,59 @@ def _record_verified_target(plan_run_root: Path, stage: int, key: str,
|
|
|
357
522
|
if (latest != supplied or final_report_data_path(report_arg.resolve()) != latest
|
|
358
523
|
or manifest.get("taskKey") != key or verified_data != data):
|
|
359
524
|
raise VerificationEvidenceError("report does not match latest verification of this task/stage")
|
|
525
|
+
captured = data["finalVerification"]["sourceImplementationReport"]["capturedHeadSha"]
|
|
526
|
+
extra: Dict[str, Any] = {}
|
|
527
|
+
head_commit = captured
|
|
528
|
+
if scope == WHOLE_TASK_SCOPE:
|
|
529
|
+
head_commit = _whole_task_stage_head(
|
|
530
|
+
rows, stage, captured,
|
|
531
|
+
Path(str(manifest.get("projectRoot") or "")), data)
|
|
532
|
+
extra = {"verification_scope": WHOLE_TASK_SCOPE,
|
|
533
|
+
"verified_head_commit": captured}
|
|
360
534
|
consumers.append_verified(
|
|
361
535
|
Path(plan_run_root), impl_task_key=key, stage=stage, verdict=verdict_token(data),
|
|
362
536
|
report_path=str(latest),
|
|
363
|
-
head_commit=
|
|
537
|
+
head_commit=head_commit,
|
|
364
538
|
final_verdict=data["finalVerdict"],
|
|
539
|
+
**extra,
|
|
365
540
|
)
|
|
366
541
|
except ValueError as exc:
|
|
367
542
|
raise HandoffError(str(exc)) from exc
|
|
368
543
|
|
|
369
544
|
|
|
545
|
+
def _whole_task_stage_head(rows: List[Dict[str, Any]], stage: int, captured: str,
|
|
546
|
+
project_root: Path, data: dict) -> str:
|
|
547
|
+
"""전체 task 검증이 이 stage 를 담았는지 확인하고, 원장에 적을 커밋을 낸다.
|
|
548
|
+
|
|
549
|
+
`stageReports` 는 보고서 자신의 주장이다. 그 stage 가 완료로 기록한 커밋이
|
|
550
|
+
검증된 head 안에 있는지까지 봐야 사실이 된다 — 검증 뒤에 갈라진 커밋은
|
|
551
|
+
검증되지 않았고, 조상 관계가 그것을 가른다.
|
|
552
|
+
|
|
553
|
+
원장에 적는 것은 검증된 head 가 아니라 이 stage 자신의 완료 커밋이다. 그
|
|
554
|
+
값이 `append_verified` 가 요구하는 "완료된 구현과 같은 커밋" 이고, PR head
|
|
555
|
+
도 그 커밋이다. 검증된 head 는 `verified_head_commit` 으로 따로 남는다.
|
|
556
|
+
"""
|
|
557
|
+
reported = (data.get("finalVerification") or {}).get("stageReports") or ()
|
|
558
|
+
covered = sorted(
|
|
559
|
+
row.get("stage") for row in reported
|
|
560
|
+
if isinstance(row, dict) and isinstance(row.get("stage"), int)
|
|
561
|
+
)
|
|
562
|
+
if stage not in covered:
|
|
563
|
+
raise HandoffError(
|
|
564
|
+
f"whole-task verification does not report stage {stage} — "
|
|
565
|
+
f"it covers {covered}")
|
|
566
|
+
head = (consumers.latest_done_by_stage(rows).get(stage) or {}).get("head_commit") or ""
|
|
567
|
+
if not head:
|
|
568
|
+
raise HandoffError(
|
|
569
|
+
f"stage {stage} has no completed implementation commit to verify")
|
|
570
|
+
if head != captured and not is_ancestor(project_root, head, captured):
|
|
571
|
+
raise HandoffError(
|
|
572
|
+
f"stage {stage} completed at {head}, which the verified head "
|
|
573
|
+
f"{captured} does not contain — re-run final-verification on a "
|
|
574
|
+
"target that includes this stage")
|
|
575
|
+
return head
|
|
576
|
+
|
|
577
|
+
|
|
370
578
|
def record_pr(*, plan_run_root, stage: int, branch: str, base: str,
|
|
371
579
|
url: str) -> Dict[str, Any]:
|
|
372
580
|
"""한 stage 의 PR 기록. PR 은 stage 당 하나이므로 행도 stage 당 하나다."""
|
|
@@ -396,7 +604,7 @@ Usage:
|
|
|
396
604
|
--project-root <dir> --project-id <id> --task-group <g> --task-id <t> \
|
|
397
605
|
--work-category <c> --stages 2,3 --base <branch>
|
|
398
606
|
okstra handoff record-verified --plan-run-root <dir> --stage <N> \
|
|
399
|
-
--report-path <
|
|
607
|
+
--report-path <report> --data-json <json>
|
|
400
608
|
okstra handoff record-pr --plan-run-root <dir> --stage <N> \
|
|
401
609
|
--branch <b> --base <b> --url <u>
|
|
402
610
|
okstra handoff local-checkout --project-root <dir> --project-id <id> \
|
|
@@ -64,17 +64,36 @@ def require_finished_verification(manifest: dict, data: dict) -> None:
|
|
|
64
64
|
|
|
65
65
|
|
|
66
66
|
def verification_row_is_current(row: dict) -> bool:
|
|
67
|
-
"""저장한 승인 행이
|
|
67
|
+
"""저장한 승인 행이 그 단계의 최신 완료 검증을 가리키는지 확인한다.
|
|
68
|
+
|
|
69
|
+
근거는 두 모양이다. 단독-stage 행은 그 stage 범위 run 을 가리키고 검증된
|
|
70
|
+
head 가 곧 그 stage 의 커밋이다. 전체 task 행은 task 범위 run 을 가리키고,
|
|
71
|
+
검증된 head 는 tip stage 의 커밋이므로 행의 `verified_head_commit` 과
|
|
72
|
+
대조한다 — 그 stage 가 보고서의 `stageReports` 에 실려 있어야 한다.
|
|
73
|
+
"""
|
|
68
74
|
try:
|
|
69
75
|
report = final_report_data_path(Path(str(row.get("report_path") or "")))
|
|
70
76
|
ref = RunRef.from_report_path(report.resolve())
|
|
71
77
|
latest, manifest, data = load_latest_verification(ref.run_dir)
|
|
72
78
|
require_finished_verification(manifest, data)
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
79
|
+
if (latest != report.resolve()
|
|
80
|
+
or manifest.get("taskKey") != row.get("impl_task_key")
|
|
81
|
+
or data.get("finalVerdict") != row.get("final_verdict")):
|
|
82
|
+
return False
|
|
83
|
+
captured = (data.get("finalVerification") or {}).get(
|
|
84
|
+
"sourceImplementationReport", {}).get("capturedHeadSha")
|
|
85
|
+
if row.get("verification_scope") == "whole-task":
|
|
86
|
+
covered = {
|
|
87
|
+
entry.get("stage")
|
|
88
|
+
for entry in (data.get("finalVerification") or {}).get("stageReports") or ()
|
|
89
|
+
if isinstance(entry, dict)
|
|
90
|
+
}
|
|
91
|
+
return (ref.stage is None
|
|
92
|
+
and data.get("verificationScope") == "whole-task"
|
|
93
|
+
and row.get("stage") in covered
|
|
94
|
+
and captured == row.get("verified_head_commit"))
|
|
95
|
+
return (ref.stage == row.get("stage")
|
|
96
|
+
and captured == row.get("head_commit"))
|
|
78
97
|
except (ValueError, JsonBoundaryError):
|
|
79
98
|
# expected-miss: 이전 형식·재검증 중인 행에는 현재 인계 자격이 없다.
|
|
80
99
|
return False
|
|
@@ -387,7 +387,16 @@ def _carry_unchanged_checklists(
|
|
|
387
387
|
def _recompute_dispatch_queue(
|
|
388
388
|
narrative: Mapping[str, Any],
|
|
389
389
|
state_items: dict[str, dict],
|
|
390
|
+
stage_ledger: Mapping[str, Any] | None = None,
|
|
390
391
|
) -> list[str]:
|
|
392
|
+
"""이월 뒤 남은 디스패치 큐.
|
|
393
|
+
|
|
394
|
+
stage 상태는 디스크 원장이 답한다. 그것 없이 서사의 depends-on 만으로 다시
|
|
395
|
+
세면 이미 구현된 stage 가 `ready` 로 읽혀 큐에 들어가고, 정작 이번 run 이
|
|
396
|
+
추가한 stage 는 의존이 안 풀린 것으로 보여 빠진다 — 큐가 정확히 뒤집힌다
|
|
397
|
+
(2026-09-24, jobs implementation-planning 003: stage 1~8 done + stage 9
|
|
398
|
+
추가인 carry-all run 에서 큐 33 → 60, 내용은 done stage 의 항목뿐).
|
|
399
|
+
"""
|
|
391
400
|
try:
|
|
392
401
|
extracted = extract_plan_items(_planning(dict(narrative)))
|
|
393
402
|
except (CarryError, PlanItemContractError, KeyError, TypeError):
|
|
@@ -404,7 +413,13 @@ def _recompute_dispatch_queue(
|
|
|
404
413
|
}
|
|
405
414
|
if not previous_hashes:
|
|
406
415
|
return []
|
|
407
|
-
ledger = planning_stage_ledger(
|
|
416
|
+
ledger = planning_stage_ledger(
|
|
417
|
+
_planning(dict(narrative)),
|
|
418
|
+
{
|
|
419
|
+
str(stage): str(status)
|
|
420
|
+
for stage, status in (stage_ledger or {}).items()
|
|
421
|
+
},
|
|
422
|
+
)
|
|
408
423
|
return reverify_item_ids(extracted, previous_hashes, ledger)
|
|
409
424
|
|
|
410
425
|
|
|
@@ -479,7 +494,9 @@ def merge_v3_plan_state(
|
|
|
479
494
|
pbv["planItems"] = [
|
|
480
495
|
state_items[item_id] for item_id in sorted(state_items)
|
|
481
496
|
]
|
|
482
|
-
queue = _recompute_dispatch_queue(
|
|
497
|
+
queue = _recompute_dispatch_queue(
|
|
498
|
+
narrative, state_items, pbv.get("stageLedger"),
|
|
499
|
+
)
|
|
483
500
|
if queue:
|
|
484
501
|
pbv["dispatchQueue"] = queue
|
|
485
502
|
return state
|
|
@@ -1579,6 +1579,13 @@ def _reject_uncompleted_round_loss(
|
|
|
1579
1579
|
|
|
1580
1580
|
이력이 없는 데이터(final-report `--data` 경로)는 대상이 아니다 — 라운드
|
|
1581
1581
|
스냅샷을 갖는 것은 convergence 소유 상태 파일뿐이다.
|
|
1582
|
+
|
|
1583
|
+
직전 run 에서 이월된 표도 대상이 아니다. 그 행의 라운드 번호는 **그 run 의**
|
|
1584
|
+
번호라 이번 run 의 이력에는 없고, 그래서 열린 라운드로 읽혔다 — 안내하는
|
|
1585
|
+
`complete-round --round 2` 는 이번 run 에 존재하지도 않는 라운드라 실행할 수
|
|
1586
|
+
없었다(2026-09-24, jobs implementation-planning 003: seed --prior-state 가
|
|
1587
|
+
이월한 P-Dep·P-Var 5건이 round 1 적용을 막음). 그 표의 이력은 직전 run 의
|
|
1588
|
+
상태 파일이 갖고 있다.
|
|
1582
1589
|
"""
|
|
1583
1590
|
if not isinstance(history, list):
|
|
1584
1591
|
return
|
|
@@ -1594,6 +1601,8 @@ def _reject_uncompleted_round_loss(
|
|
|
1594
1601
|
for row in item.get("verdicts") or []:
|
|
1595
1602
|
if not isinstance(row, Mapping):
|
|
1596
1603
|
continue
|
|
1604
|
+
if str(row.get("carriedForwardFromSeq") or "").strip():
|
|
1605
|
+
continue
|
|
1597
1606
|
recorded_round = row.get("round")
|
|
1598
1607
|
if (
|
|
1599
1608
|
isinstance(recorded_round, int)
|
|
@@ -2010,6 +2010,35 @@ def _initialize_report_ledgers(ctx: Mapping[str, Any], manifest: Mapping[str, An
|
|
|
2010
2010
|
activity_path = Path(activity_value)
|
|
2011
2011
|
if activity_value and not activity_path.is_file():
|
|
2012
2012
|
_write_text(activity_path, "")
|
|
2013
|
+
_initialize_lead_only_convergence(ctx, manifest)
|
|
2014
|
+
|
|
2015
|
+
|
|
2016
|
+
def _initialize_lead_only_convergence(
|
|
2017
|
+
ctx: Mapping[str, Any], manifest: Mapping[str, Any],
|
|
2018
|
+
) -> None:
|
|
2019
|
+
"""워커를 띄우지 않는 phase 의 수렴 상태.
|
|
2020
|
+
|
|
2021
|
+
리포트 조립은 모든 리포트에 수렴 입력을 요구하는데(ADR-0020), 프로필이
|
|
2022
|
+
"convergence 를 돌지 않는다" 고 선언한 phase 에는 그 파일을 만드는 주체가
|
|
2023
|
+
없었다. 리드가 매 run 마다 seed → plan-round → finalize 를 손으로 돌려 빈
|
|
2024
|
+
상태를 만들어 왔다(실측: dev-10860 release-handoff 001·002 둘 다). 나머지
|
|
2025
|
+
입력과 같은 자리에서 같은 조건으로 만든다 — 없을 때만.
|
|
2026
|
+
"""
|
|
2027
|
+
from .convergence_engine import lead_only_final_state
|
|
2028
|
+
from .convergence_store import write_final_state_atomic
|
|
2029
|
+
|
|
2030
|
+
if _resolve_workers(dict(ctx)):
|
|
2031
|
+
return
|
|
2032
|
+
value = str(ctx.get("CONVERGENCE_STATE_PATH") or "")
|
|
2033
|
+
path = Path(value)
|
|
2034
|
+
if not value or path.exists():
|
|
2035
|
+
return
|
|
2036
|
+
path.parent.mkdir(parents=True, exist_ok=True)
|
|
2037
|
+
write_final_state_atomic(
|
|
2038
|
+
path,
|
|
2039
|
+
lead_only_final_state(str(manifest.get("taskKey") or "")),
|
|
2040
|
+
migration=None,
|
|
2041
|
+
)
|
|
2013
2042
|
|
|
2014
2043
|
|
|
2015
2044
|
_MODEL_ASSIGNMENT_KEYS = {
|
|
@@ -728,6 +728,12 @@ def _handoff_eligibility(state: WizardState) -> list:
|
|
|
728
728
|
return compute_eligibility(stage_map, rows)
|
|
729
729
|
|
|
730
730
|
|
|
731
|
+
# 자격 stage 전체를 한 번에 고르는 값. stage 번호는 숫자뿐이라 충돌하지 않는다.
|
|
732
|
+
# CLI 는 `--stages` 를 비우면 전체를 쓰지만, 위저드 pick 은 빈 선택을 거절하므로
|
|
733
|
+
# 그 뜻을 표현할 값이 따로 필요하다.
|
|
734
|
+
ALL_ELIGIBLE_STAGES = "all"
|
|
735
|
+
|
|
736
|
+
|
|
731
737
|
def _build_handoff_stage_pick(state: WizardState) -> Prompt:
|
|
732
738
|
elig = _handoff_eligibility(state)
|
|
733
739
|
eligible = [e for e in elig if e["eligible"]]
|
|
@@ -742,6 +748,9 @@ def _build_handoff_stage_pick(state: WizardState) -> Prompt:
|
|
|
742
748
|
t = _p(state.workspace_root, "handoff_stage_pick", blocked=blocked_summary)
|
|
743
749
|
stage_label = t["labels"]["stage"]
|
|
744
750
|
options: list[Option] = []
|
|
751
|
+
if len(eligible) > 1:
|
|
752
|
+
options.append(_opt(ALL_ELIGIBLE_STAGES, t["labels"]["all_stages"].format(
|
|
753
|
+
stages=", ".join(str(e["stage"]) for e in eligible))))
|
|
745
754
|
for e in eligible:
|
|
746
755
|
deps = ", ".join(str(d) for d in e["depends_on"]) or "-"
|
|
747
756
|
options.append(_opt(str(e["stage"]),
|
|
@@ -760,6 +769,8 @@ def _submit_handoff_stage_pick(state: WizardState, value: str) -> Optional[str]:
|
|
|
760
769
|
raise WizardError(t["errors"]["none_selected"])
|
|
761
770
|
eligible = {str(e["stage"]) for e in _handoff_eligibility(state)
|
|
762
771
|
if e["eligible"]}
|
|
772
|
+
if ALL_ELIGIBLE_STAGES in picks:
|
|
773
|
+
picks = sorted(eligible, key=int)
|
|
763
774
|
bad = [p for p in picks if p not in eligible]
|
|
764
775
|
if bad:
|
|
765
776
|
raise WizardError(t["errors"]["not_eligible"].format(
|
|
@@ -1429,6 +1429,29 @@ def _dispatch_roster_key(row: Mapping[str, Any]) -> str:
|
|
|
1429
1429
|
return ""
|
|
1430
1430
|
|
|
1431
1431
|
|
|
1432
|
+
def _dispatch_row_paths(team_state: Mapping[str, Any], worker_id: str) -> tuple[str, str]:
|
|
1433
|
+
"""이 로스터 워커의 dispatch 행이 기록한 (promptPath, resultPath).
|
|
1434
|
+
|
|
1435
|
+
v2 assignment 로 띄운 워커는 로스터 행의 경로가 빈 채로 남는다.
|
|
1436
|
+
`workers[].promptPath` 는 첫 v1 dispatch 하나만 가리키는 필드이고 v2 행은
|
|
1437
|
+
그것을 채우지 않는다(`okstra_ctl.dispatch_state.worker_dispatch_records` 의
|
|
1438
|
+
주석, 2026-09-08 실측). 그래서 acceptance critic 처럼 v2 로만 띄우는 워커는
|
|
1439
|
+
프롬프트와 결과 파일이 디스크에 그대로 있는데도 "promptPath 가 없다",
|
|
1440
|
+
"결과 파일이 없다" 로 보고됐다(2026-09-24, jobs final-verification 002).
|
|
1441
|
+
사실은 dispatch 행에 있으므로 그리로 폴백한다.
|
|
1442
|
+
"""
|
|
1443
|
+
prompt = ""
|
|
1444
|
+
result = ""
|
|
1445
|
+
for row in team_state.get("workerDispatches") or ():
|
|
1446
|
+
if not isinstance(row, Mapping) or _dispatch_roster_key(row) != worker_id:
|
|
1447
|
+
continue
|
|
1448
|
+
prompt = prompt or str(row.get("promptPath") or "")
|
|
1449
|
+
result = result or str(
|
|
1450
|
+
row.get("workerResultPath") or row.get("resultPath") or ""
|
|
1451
|
+
)
|
|
1452
|
+
return prompt, result
|
|
1453
|
+
|
|
1454
|
+
|
|
1432
1455
|
def _validate_cmux_workers_were_dispatched_by_okstra(
|
|
1433
1456
|
team_state: dict,
|
|
1434
1457
|
workers: list,
|
|
@@ -1628,14 +1651,26 @@ def validate_team_state(
|
|
|
1628
1651
|
f"{role} must use modelExecutionValue `{expected_model_execution_value}`"
|
|
1629
1652
|
)
|
|
1630
1653
|
|
|
1654
|
+
# 경로는 로스터 행이 소유하지만 v2 dispatch 는 그 행을 채우지 않는다.
|
|
1655
|
+
# 비어 있을 때만 dispatch 행에서 읽는다 — 로스터 행에 값이 있으면 그것이
|
|
1656
|
+
# 대조 대상이다.
|
|
1657
|
+
roster_prompt = str(worker.get("promptPath") or "")
|
|
1658
|
+
roster_result = str(worker.get("resultPath") or "")
|
|
1659
|
+
if not roster_prompt or not roster_result:
|
|
1660
|
+
fallback_prompt, fallback_result = _dispatch_row_paths(
|
|
1661
|
+
team_state, str(worker.get("workerId") or "")
|
|
1662
|
+
)
|
|
1663
|
+
roster_prompt = roster_prompt or fallback_prompt
|
|
1664
|
+
roster_result = roster_result or fallback_result
|
|
1665
|
+
|
|
1631
1666
|
expected_result_relative = expected.get("resultPath")
|
|
1632
|
-
result_relative =
|
|
1667
|
+
result_relative = roster_result
|
|
1633
1668
|
if expected_result_relative and result_relative != expected_result_relative:
|
|
1634
1669
|
failures.append(
|
|
1635
1670
|
f"{role} must use resultPath `{expected_result_relative}`"
|
|
1636
1671
|
)
|
|
1637
1672
|
expected_prompt_relative = expected.get("promptPath")
|
|
1638
|
-
prompt_relative =
|
|
1673
|
+
prompt_relative = roster_prompt
|
|
1639
1674
|
if expected_prompt_relative and prompt_relative != expected_prompt_relative:
|
|
1640
1675
|
failures.append(
|
|
1641
1676
|
f"{role} must use promptPath `{expected_prompt_relative}`"
|
|
@@ -5778,7 +5813,7 @@ def _validate_verified_row_recorded(
|
|
|
5778
5813
|
report_path: Path,
|
|
5779
5814
|
failures: list[str],
|
|
5780
5815
|
) -> None:
|
|
5781
|
-
"""A release-ready
|
|
5816
|
+
"""A release-ready verification must leave a `verified` row per cleared stage.
|
|
5782
5817
|
|
|
5783
5818
|
`okstra handoff record-verified` validates its own inputs, but nothing
|
|
5784
5819
|
checked that it ever ran. Skipping it leaves the report saying `accepted`
|
|
@@ -5787,7 +5822,8 @@ def _validate_verified_row_recorded(
|
|
|
5787
5822
|
the only recoveries are re-running an expensive phase or hand-editing the
|
|
5788
5823
|
registry.
|
|
5789
5824
|
"""
|
|
5790
|
-
|
|
5825
|
+
scope = str(data.get("verificationScope") or "")
|
|
5826
|
+
if scope not in ("single-stage", "whole-task"):
|
|
5791
5827
|
return
|
|
5792
5828
|
if not release_handoff_allowed(data):
|
|
5793
5829
|
return
|
|
@@ -5802,11 +5838,17 @@ def _validate_verified_row_recorded(
|
|
|
5802
5838
|
if rows is None:
|
|
5803
5839
|
return
|
|
5804
5840
|
head = ((data.get("finalVerification") or {}).get("sourceImplementationReport") or {}).get("capturedHeadSha")
|
|
5841
|
+
# 전체 task 행의 `head_commit` 은 그 stage 자신의 완료 커밋이다. 검증된 head
|
|
5842
|
+
# 는 `verified_head_commit` 에 있으므로, 그 행을 stage 커밋으로 찾으면 모든
|
|
5843
|
+
# stage 가 누락으로 읽힌다.
|
|
5844
|
+
head_field = (
|
|
5845
|
+
"verified_head_commit" if scope == "whole-task" else "head_commit"
|
|
5846
|
+
)
|
|
5805
5847
|
verified = {
|
|
5806
5848
|
row.get("stage")
|
|
5807
5849
|
for row in rows
|
|
5808
5850
|
if isinstance(row, dict) and row.get("status") == "verified"
|
|
5809
|
-
and head and row.get(
|
|
5851
|
+
and head and row.get(head_field) == head
|
|
5810
5852
|
and row.get("final_verdict") == data.get("finalVerdict")
|
|
5811
5853
|
and _data_path_for(Path(str(row.get("report_path") or ""))).resolve() == _data_path_for(report_path).resolve()
|
|
5812
5854
|
}
|
|
@@ -5816,9 +5858,9 @@ def _validate_verified_row_recorded(
|
|
|
5816
5858
|
f"final-verification cleared stage(s) {missing} for release but "
|
|
5817
5859
|
"`runs/implementation-planning/consumers.jsonl` carries no "
|
|
5818
5860
|
"`verified` row matching this report and captured commit. Run `okstra handoff record-verified` "
|
|
5819
|
-
"before finishing — without that row the
|
|
5820
|
-
"while the registry says unverified, and the
|
|
5821
|
-
"offered for a
|
|
5861
|
+
"before finishing — once per cleared stage — without that row the "
|
|
5862
|
+
"report says accepted while the registry says unverified, and the "
|
|
5863
|
+
"stage is never offered for a pull request "
|
|
5822
5864
|
'(final-verification.md §"Verified-row recording").'
|
|
5823
5865
|
)
|
|
5824
5866
|
|
|
@@ -383,6 +383,9 @@ def _collect_lead_evidence(
|
|
|
383
383
|
if candidate is not None:
|
|
384
384
|
sessions.setdefault(sid, candidate)
|
|
385
385
|
evidence = _LeadEvidence(window=(since, until))
|
|
386
|
+
ledger_progress = _ledger_progress(
|
|
387
|
+
team_state, run_manifest, project_root, task_type, suffix
|
|
388
|
+
)
|
|
386
389
|
for sid, path in sorted(sessions.items()):
|
|
387
390
|
progress, reads, agent_name = _scan_one_jsonl(path, since, until)
|
|
388
391
|
if agent_name and sid != lead_sid:
|
|
@@ -400,7 +403,10 @@ def _collect_lead_evidence(
|
|
|
400
403
|
"implementation entry-guard conformance cannot be verified, which "
|
|
401
404
|
"fails the run (same principle as the token-usage accuracy contract)."
|
|
402
405
|
)
|
|
403
|
-
|
|
406
|
+
# 전사와 원장은 같은 체크포인트의 두 기록이다. 계약은 둘 다 요구하고
|
|
407
|
+
# (`lead-progress append` 로 기록, 같은 줄을 대화에 raw 로 emit), 검사는
|
|
408
|
+
# 어느 쪽에 남았든 그 체크포인트를 본 것으로 판정한다.
|
|
409
|
+
evidence.progress = _merge_progress(evidence.progress, ledger_progress)
|
|
404
410
|
for ts_list in evidence.sidecar_reads.values():
|
|
405
411
|
ts_list.sort()
|
|
406
412
|
if _is_activity_contract_v1_planning(run_manifest):
|
|
@@ -643,6 +649,69 @@ def _collect_artifact_lead_evidence(
|
|
|
643
649
|
return evidence, None
|
|
644
650
|
|
|
645
651
|
|
|
652
|
+
def _ledger_progress(
|
|
653
|
+
team_state: dict,
|
|
654
|
+
run_manifest: Mapping[str, Any],
|
|
655
|
+
project_root: Path,
|
|
656
|
+
task_type: str,
|
|
657
|
+
suffix: str | None,
|
|
658
|
+
) -> list[tuple[str, str, str]]:
|
|
659
|
+
"""원장에 기록된 이 run 의 PROGRESS 체크포인트. 못 읽으면 빈 목록.
|
|
660
|
+
|
|
661
|
+
`okstra lead-progress append` 는 호스트와 무관하게 체크포인트를
|
|
662
|
+
`leadEventsPath` 에 쓰고, 리드 계약은 모든 체크포인트를 그 명령으로
|
|
663
|
+
기록하라고 요구한다(`prompts/lead/okstra-lead-contract.md` "Progress
|
|
664
|
+
reporting"). 그런데 `claude-jsonl` 증거 경로는 세션 전사만 훑어서, 계약대로
|
|
665
|
+
기록한 run 이 체크포인트 전건 누락으로 보고됐다 — 기본 호스트에서 그 명령의
|
|
666
|
+
출력을 읽는 소비자가 없었다(2026-09-24, jobs final-verification 002:
|
|
667
|
+
원장에 progress 30행, advisory 10건).
|
|
668
|
+
|
|
669
|
+
원장 행은 스크랩한 대화 텍스트보다 약한 증거가 아니다. `--phase` 는 열거된
|
|
670
|
+
phase id 만 받고 `--worker` 는 로스터 역할로 다시 쓰이므로, 그 행은 검증된
|
|
671
|
+
입력으로 okstra 자신이 쓴 것이다.
|
|
672
|
+
"""
|
|
673
|
+
events_path, _error = _resolve_lead_events_path(
|
|
674
|
+
team_state, run_manifest, project_root
|
|
675
|
+
)
|
|
676
|
+
if events_path is None:
|
|
677
|
+
return []
|
|
678
|
+
try:
|
|
679
|
+
events = read_lead_events(events_path)
|
|
680
|
+
except LeadEventParseError:
|
|
681
|
+
return []
|
|
682
|
+
run_seq = _run_sequence(run_manifest, suffix)
|
|
683
|
+
rows: list[tuple[str, str, str]] = []
|
|
684
|
+
for event in events:
|
|
685
|
+
if event.event_type not in ("progress", "progress-checkpoint"):
|
|
686
|
+
continue
|
|
687
|
+
if not _event_matches_run(event, team_state, run_manifest, task_type, run_seq):
|
|
688
|
+
continue
|
|
689
|
+
progress = _progress_line_from_event(event)
|
|
690
|
+
if progress is not None:
|
|
691
|
+
rows.append(progress)
|
|
692
|
+
return rows
|
|
693
|
+
|
|
694
|
+
|
|
695
|
+
def _merge_progress(
|
|
696
|
+
rows: list[tuple[str, str, str]],
|
|
697
|
+
extra: list[tuple[str, str, str]],
|
|
698
|
+
) -> list[tuple[str, str, str]]:
|
|
699
|
+
"""두 증거 출처의 체크포인트를 합친다. 같은 줄은 한 번만 남는다.
|
|
700
|
+
|
|
701
|
+
리드는 계약상 원장에 기록하고 같은 줄을 대화에 내보내므로, 합치면 같은
|
|
702
|
+
체크포인트가 두 번 들어온다. 서술 정확성 검사는 줄 단위로 대조하니 중복은
|
|
703
|
+
판정을 바꾸지 않지만, 보고 문구에 같은 줄이 두 번 실리는 것을 막는다.
|
|
704
|
+
"""
|
|
705
|
+
seen = {(phase, line) for _ts, phase, line in rows}
|
|
706
|
+
for ts, phase, line in extra:
|
|
707
|
+
if (phase, line) in seen:
|
|
708
|
+
continue
|
|
709
|
+
seen.add((phase, line))
|
|
710
|
+
rows.append((ts, phase, line))
|
|
711
|
+
rows.sort()
|
|
712
|
+
return rows
|
|
713
|
+
|
|
714
|
+
|
|
646
715
|
def _conformance_evidence_source(
|
|
647
716
|
team_state: dict,
|
|
648
717
|
) -> tuple[str | None, str | None]:
|