sortie-dogs 0.13.11 → 0.13.13
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +157 -650
- package/dist/asset-version.d.ts +1 -1
- package/dist/asset-version.js +1 -1
- package/dist/core/mission-context.d.ts +73 -0
- package/dist/core/mission-context.js +66 -0
- package/dist/core/mission-validation-command.d.ts +2 -0
- package/dist/core/mission-validation-command.js +10 -0
- package/dist/core/operator-mission.d.ts +8 -2
- package/dist/core/operator-mission.js +46 -10
- package/dist/core/operator-runtime.d.ts +2 -0
- package/dist/core/operator-runtime.js +28 -0
- package/dist/core/run-flight-ledger.d.ts +1 -0
- package/dist/core/run-flight-ledger.js +5 -2
- package/dist/index.d.ts +1 -1
- package/dist/index.js +1 -1
- package/dist/plugin/gate.js +16 -1
- package/dist/plugin/profiled.js +73 -17
- package/dist/plugin/protected-snapshot.js +10 -4
- package/dist/plugin/v2.js +4 -4
- package/dist/runtime-mission-assets.d.ts +4 -2
- package/dist/runtime-mission-assets.js +25 -9
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -4,729 +4,236 @@
|
|
|
4
4
|
<img src="docs/assets/sortie-dogs-logo.png" alt="Sortie-dogs logo" width="640">
|
|
5
5
|
</p>
|
|
6
6
|
|
|
7
|
-
**
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
Use
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
- **Goal invariance**: accepted outcomes and proof requirements survive delegation,
|
|
14
|
-
continuation, remediation, and restart.
|
|
15
|
-
- **Adaptive execution**: small work stays small; additional agents and stronger
|
|
16
|
-
models are used only when task shape or risk justifies them.
|
|
17
|
-
- **Clear Worker instructions**: give Luna a concise goal, explicit constraints
|
|
18
|
-
and completion criteria; keep orchestration bookkeeping in the harness.
|
|
19
|
-
See the [instruction design principles](docs/worker-instruction-design.md).
|
|
20
|
-
- **Coexistence**: Sortie activates only when selected and preserves normal
|
|
21
|
-
OpenCode agents, settings, and user-owned files.
|
|
22
|
-
- **Cost, time, and proof**: the objective is a verified result at the lowest
|
|
23
|
-
practical cost and wall time, not the largest agent count.
|
|
7
|
+
**An adaptive execution harness for OpenCode and Codex: preserve the goal,
|
|
8
|
+
implement, validate, review, and deliver without taking over your setup.**
|
|
9
|
+
|
|
10
|
+
Use your host normally. Invoke Sortie when a task needs coordinated execution and an
|
|
11
|
+
evidence-backed result, not more agents for their own sake.
|
|
24
12
|
|
|
25
13
|
[](https://github.com/zufall-upon/Sortie-dogs/releases/latest)
|
|
26
14
|
[](https://www.npmjs.com/package/sortie-dogs)
|
|
27
15
|
[](https://github.com/zufall-upon/Sortie-dogs/actions/workflows/test.yml)
|
|
28
|
-
[](https://opencode.ai/)
|
|
17
|
+
[](docs/codex.md)
|
|
29
18
|
[](https://www.typescriptlang.org/)
|
|
30
19
|
[](https://www.npmjs.com/package/sortie-dogs)
|
|
31
20
|
[](LICENSE)
|
|
32
21
|
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
Guides: [日本語](docs/guide-ja.md) · [简体中文](docs/guide-zh-CN.md) ·
|
|
36
|
-
[Testing](docs/testing.md) · [CLI testing](docs/cli-testing.md)
|
|
37
|
-
|
|
38
|
-
**Current release: [v0.13.10](https://github.com/zufall-upon/Sortie-dogs/releases/tag/v0.13.10)**
|
|
39
|
-
([release notes](docs/release-v0.13.10.md)). The default Mission runtime retains the `v010`
|
|
40
|
-
profile, command and configuration names for compatibility; these names do not mean v0.10 is installed.
|
|
41
|
-
The current asset marker is `0.13.10-codex-windows-v1`.
|
|
42
|
-
|
|
43
|
-
## SWE-bench Lite: 170/300 (56.67%)
|
|
44
|
-
|
|
45
|
-
The fixed **Sortie-dogs v0.12.24** harness resolved **170 of 300 SWE-bench Lite test issues** in one pass@1 campaign, with 9 empty patches and no official evaluation errors. Every instance has a frozen prediction and an inference-time trajectory. The task Workers ran `openai/gpt-6-luna-fast#max`; operator, coordinator and review roles ran `openai/gpt-6-sol#xhigh`. This is a system result, **not** a Luna-only model comparison or a Verified/full SWE-bench score.
|
|
22
|
+
**Current release: [v0.13.13](https://github.com/zufall-upon/Sortie-dogs/releases/tag/v0.13.13)**
|
|
23
|
+
· [Release notes](docs/release-v0.13.13.md)
|
|
46
24
|
|
|
47
|
-
[
|
|
25
|
+
[Quick start](#quick-start) · [Workflow](#mission-workflow) · [Models](#default-models) ·
|
|
26
|
+
[Measured results](#measured-results) · [Configuration](#configuration) · [Documentation](#documentation)
|
|
48
27
|
|
|
49
|
-
|
|
28
|
+
> **Beta:** the v0.13.x package is still stabilizing before 1.0. Both host integrations are
|
|
29
|
+
> implemented; Codex Missions have completed implementation, validation, correction and acceptance.
|
|
30
|
+
> [Host-specific execution boundaries](docs/codex.md#native-permissions-and-approval) remain explicit.
|
|
50
31
|
|
|
51
|
-
|
|
52
|
-
optional measurement rather than a mandatory release gate.
|
|
32
|
+

|
|
53
33
|
|
|
54
|
-
|
|
55
|
-
|
|
34
|
+
## Why Sortie
|
|
35
|
+
|
|
36
|
+
- **Autonomy:** retain the original request across delegation, correction, compaction and restart.
|
|
37
|
+
Finish in-request work through the existing host authority instead of routine approval round trips.
|
|
38
|
+
- **Efficiency:** small work stays small. Use a direct Worker for a known single unit; add
|
|
39
|
+
Coordinator, Scout or Advisor only when discovery or task shape calls for them. Reuse valid
|
|
40
|
+
unchanged evidence instead of repeating successful work.
|
|
41
|
+
- **Visibility:** show real commands, exits, timing, model routes, review disposition and final
|
|
42
|
+
acceptance. Unknown execution or cost stays unknown, not an invented success or zero.
|
|
43
|
+
- **Quality:** meaningful formal checks and risk-based review establish completion, not model prose.
|
|
44
|
+
Worker instructions carry a concise goal, explicit constraints and completion criteria;
|
|
45
|
+
the harness owns bookkeeping. [Instruction design](docs/worker-instruction-design.md).
|
|
46
|
+
- **Coexistence:** explicit invocation, existing authentication and host settings, one shared
|
|
47
|
+
Mission lifecycle. Normal OpenCode/Codex use and user-owned files remain yours.
|
|
56
48
|
|
|
57
49
|
## Quick start
|
|
58
50
|
|
|
59
|
-
|
|
51
|
+
Choose **one host route**. Both require Node.js **22.6+** and npm; project-local installation is
|
|
52
|
+
recommended. OpenCode is not required for Codex, and Codex is not required for OpenCode.
|
|
60
53
|
|
|
61
|
-
|
|
62
|
-
Use Node.js 22.6 or newer and an existing Codex CLI with ChatGPT authentication. Ubuntu validation
|
|
63
|
-
used Node.js 22.22.1 and Codex 0.160.1. Sortie neither installs Codex nor starts a login flow, copies
|
|
64
|
-
credentials, or creates a second host configuration. Mission execution refuses non-ChatGPT auth.
|
|
54
|
+
### Codex
|
|
65
55
|
|
|
66
|
-
|
|
67
|
-
|
|
56
|
+
Use an existing Codex CLI signed in with **ChatGPT**. On Windows, PowerShell **7** must already be
|
|
57
|
+
available; POSIX uses bash. Sortie does not install Codex, start login or use a metered API fallback.
|
|
68
58
|
|
|
69
59
|
```sh
|
|
70
|
-
npm install --save-dev sortie-dogs@0.13.
|
|
60
|
+
npm install --save-dev sortie-dogs@0.13.13
|
|
71
61
|
npx --no-install sortie-dogs codex init .
|
|
72
|
-
npx --no-install sortie-dogs codex mission --help
|
|
73
62
|
```
|
|
74
63
|
|
|
75
|
-
|
|
76
|
-
can produce a local tarball with `npm pack` (build included) and install that tarball instead.
|
|
77
|
-
This does not publish it. Use the same installed version for the CLI and SDK below.
|
|
78
|
-
`codex init` installs only `.agents/skills/sortie-dogs`; it does not create or edit `.opencode`.
|
|
79
|
-
It refuses to overwrite a skill whose ownership cannot be established.
|
|
80
|
-
|
|
81
|
-
After restarting Codex or opening the project in a new chat, invoke the installed skill explicitly:
|
|
64
|
+
Restart Codex or open the project in a new chat, then invoke the installed skill explicitly:
|
|
82
65
|
|
|
83
66
|
```text
|
|
84
67
|
$sortie-dogs Implement and verify the requested change
|
|
85
68
|
```
|
|
86
69
|
|
|
87
|
-
|
|
88
|
-
as exact `/sortie-dogs` slash commands; `/skills` opens the skill picker. The skill is intentionally
|
|
89
|
-
explicit-only so child Codex Mission turns cannot invoke another Sortie Mission recursively.
|
|
90
|
-
On POSIX and Windows it selects the natural-language, multi-role Mission route below.
|
|
91
|
-
Windows uses existing PowerShell 7; it does not route natural-language Missions through the
|
|
92
|
-
single-task `codex run` manifest workflow. That workflow remains available when explicitly selected
|
|
93
|
-
and retains its post-execution observation boundary, not a pre-execution guard.
|
|
94
|
-
|
|
95
|
-
Sortie-dogs also exports a host adapter for the official Codex app-server stdio protocol. This is
|
|
96
|
-
separate from selecting an OpenAI model through OpenCode: OpenCode model routing still uses the
|
|
97
|
-
OpenCode host, while `CodexAppServerHost` starts and observes native Codex threads and turns.
|
|
98
|
-
|
|
99
|
-
```ts
|
|
100
|
-
import { CodexAppServerHost, createCodexAppServerTransport } from "sortie-dogs";
|
|
101
|
-
|
|
102
|
-
const transport = createCodexAppServerTransport({ executable: "codex" });
|
|
103
|
-
const host = new CodexAppServerHost(transport);
|
|
104
|
-
const threadID = await host.startThread({ cwd: process.cwd(), ephemeral: true });
|
|
105
|
-
const result = await host.runTurn(threadID, "Implement the admitted unit.", {
|
|
106
|
-
cwd: process.cwd(),
|
|
107
|
-
approvalPolicy: "unlessTrusted",
|
|
108
|
-
sandboxPolicy: {
|
|
109
|
-
type: "workspaceWrite",
|
|
110
|
-
writableRoots: [process.cwd()],
|
|
111
|
-
networkAccess: false,
|
|
112
|
-
},
|
|
113
|
-
});
|
|
114
|
-
await host.close();
|
|
115
|
-
```
|
|
116
|
-
|
|
117
|
-
The adapter reuses Sortie's host-neutral core instead of copying its goal, acceptance, evidence, or
|
|
118
|
-
ledger implementation. It exposes completed Codex items as authoritative observations, supports
|
|
119
|
-
exact persisted-thread resume and turn interruption (ephemeral threads have no resumable rollout), and declines command/file approvals unless the caller
|
|
120
|
-
provides an approval handler. Permission-subset requests receive an empty grant unless the caller
|
|
121
|
-
connects `permissionsApproval`; approved responses are limited to the native request. Unsupported
|
|
122
|
-
server-initiated requests fail closed.
|
|
123
|
-
|
|
124
|
-
For the existing Operator → Coordinator → Worker → Reviewer Mission workflow, use:
|
|
70
|
+
Or run the same natural-language Mission from the terminal:
|
|
125
71
|
|
|
126
72
|
```sh
|
|
127
73
|
npx --no-install sortie-dogs codex mission --prompt "Implement and review the requested change"
|
|
128
74
|
```
|
|
129
75
|
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
// trustedPowerShellExecutable: "C:/Program Files/PowerShell/7/pwsh.exe", // Windows, optional.
|
|
140
|
-
// resumeThreadID: "saved-root-thread-id", // Continue the same Mission when needed.
|
|
141
|
-
});
|
|
142
|
-
try {
|
|
143
|
-
const result = await mission.run("Implement and review the requested change");
|
|
144
|
-
console.log(JSON.stringify(result));
|
|
145
|
-
process.exitCode = result.accepted ? 0 : 1;
|
|
146
|
-
} finally {
|
|
147
|
-
await mission.close();
|
|
148
|
-
}
|
|
149
|
-
```
|
|
76
|
+
`codex init` installs `.agents/skills/sortie-dogs` only; it does not edit `.opencode`.
|
|
77
|
+
**Do not run `sortie-dogs init` for this route.** `$sortie-dogs` is Skills syntax, not an exact
|
|
78
|
+
`/sortie-dogs` slash command. Leave model overrides unset to retain the role defaults.
|
|
79
|
+
|
|
80
|
+
Existing native permissions remain authoritative. An SDK application can connect its already
|
|
81
|
+
authorized parent-host executor; the CLI does not supply one or an interactive native approval
|
|
82
|
+
bridge. [Codex guide: SDK, permissions, progress, resume and manifest route](docs/codex.md).
|
|
83
|
+
|
|
84
|
+
### OpenCode
|
|
150
85
|
|
|
151
|
-
|
|
152
|
-
validation evidence, and final Operator acceptance. It runs saved native Codex threads with existing
|
|
153
|
-
ChatGPT authentication; it does not import OpenCode settings or introduce a second Mission ledger.
|
|
154
|
-
Commands run through `/bin/bash` on POSIX or existing `pwsh.exe -NoProfile -NonInteractive -Command`
|
|
155
|
-
on Windows, using the native app-server's configured permissions. The compatibility tool remains
|
|
156
|
-
named `bash`, but its Windows input must be PowerShell syntax. Windows callers can pin an absolute
|
|
157
|
-
PowerShell 7 path with SDK `trustedPowerShellExecutable` or CLI `--trusted-pwsh`.
|
|
158
|
-
Packaged roles retain Sol 6.1/xhigh for Operator, Coordinator and Reviewer. The shared Luna-fast/max
|
|
159
|
-
Worker alias maps only in Codex to native `gpt-6-luna`/max plus `serviceTier: "priority"` (Fast).
|
|
160
|
-
OpenCode routing is unchanged. Fast uses subscription limits faster than Standard; no metered API
|
|
161
|
-
fallback is introduced. CLI progress exposes each role's native model, effort and separate service tier.
|
|
162
|
-
Sortie does not
|
|
163
|
-
replace them with a fixed repository-only or network-disabled policy, change host configuration, or select
|
|
164
|
-
full access. Thread turns retain native thread permissions; standalone `command/exec` uses the server's
|
|
165
|
-
configured policy, not a thread's temporary grants. A parent application's in-memory approval is not
|
|
166
|
-
automatically transferred to a separately launched app-server.
|
|
167
|
-
Windows recovery can reclaim a killed adapter only when its recorded Windows PID is absent.
|
|
168
|
-
A live/reused PID, access denial, legacy owner without platform identity, or unresolved native
|
|
169
|
-
execution remains unknown; no command is replayed. Sandbox startup errors must be resolved at the
|
|
170
|
-
native host, not by silently widening Sortie's permissions or stopping unrelated Codex processes.
|
|
171
|
-
SDK hosts can forward native command/file approval requests through `approval` and permission-subset
|
|
172
|
-
requests through `permissionsApproval`. The native host remains responsible for deciding and enforcing
|
|
173
|
-
the grant. The CLI reports approval requests but has no interactive approval bridge; without a connected
|
|
174
|
-
host callback, no grant is returned. `command/exec` has no thread-scoped approval API, so inheriting the
|
|
175
|
-
configured policy does not add an escalation path for these standalone commands.
|
|
176
|
-
Use SDK `permissions` or CLI `--permissions <native-profile>` only to explicitly select an existing
|
|
177
|
-
native profile for thread start/resume and standalone commands. Omission retains native defaults,
|
|
178
|
-
which may be read-only. Invalid or disallowed profiles fail through the native server without fallback;
|
|
179
|
-
no configuration file is changed. Low-level `CodexAppServerHost` callers must opt into
|
|
180
|
-
`experimentalApi: true` for profile APIs; Mission sessions already negotiate that capability. Progress reports the effective native profile, sandbox and approval
|
|
181
|
-
routing. This selection does not configure a delegated parent-host executor's own permissions.
|
|
182
|
-
See [Ubuntu acceptance evidence and remaining work](docs/codex-mission-acceptance-20261006.md).
|
|
183
|
-
|
|
184
|
-
For a parent application with an existing approval-aware executor, pass the optional SDK
|
|
185
|
-
`executeCommand` callback. It receives the exact post-hook argv, cwd, timeout, thread/turn/call identity,
|
|
186
|
-
and an AbortSignal for bash/read/write operations. The parent owns approval and execution; Sortie does
|
|
187
|
-
not infer a grant from prompt text or create another permission store. Return `completed` only after
|
|
188
|
-
execution ends, with the real integer exitCode, stdout and stderr. Approval alone is not a result.
|
|
189
|
-
`denied` and `not-started` explicitly mean nothing ran and create no validation exit. `interrupted`,
|
|
190
|
-
`unknown`, exceptions and invalid results retain unknown execution and stop resends. The host owns
|
|
191
|
-
command timeout enforcement; Sortie does not add a separate deadline to its approval interaction.
|
|
192
|
-
Closing the adapter signals cancellation; late callback results cannot establish validation, and the
|
|
193
|
-
host must reconcile any still-running process. Omission keeps native `command/exec`; a connected
|
|
194
|
-
executor never falls back to native execution after rejection or failure. Existing Mission before/after
|
|
195
|
-
hooks, command identity and validation acceptance remain shared. CLI progress identifies the selected
|
|
196
|
-
executor; the CLI itself does not attach a parent executor.
|
|
197
|
-
File writes use a short argv referencing a temporary UTF-8 payload, so large files and literal quotes,
|
|
198
|
-
newlines or NUL characters do not overflow Windows command-line limits. Staging does not write the
|
|
199
|
-
project file; the same configured executor performs that write, and the payload is removed afterward.
|
|
200
|
-
Status observations (without `confirmed_conditions`) can return during an active child Task. Condition
|
|
201
|
-
registration and other modifying tools remain serialized; status does not imply Task completion.
|
|
202
|
-
The bash tool accepts `workdir` (absolute or relative to the project root). The shared hooks and both
|
|
203
|
-
executors use that directory; progress and receipts include the effective cwd. A successful command in
|
|
204
|
-
another directory does not satisfy a validation declared for the project root.
|
|
205
|
-
|
|
206
|
-
Packaged role models and reasoning levels apply by default. `--model` and `--effort` explicitly
|
|
207
|
-
override all roles; SDK callers can use `roleModels` for individual roles. Unsupported native model
|
|
208
|
-
names fail rather than silently selecting another model. Availability depends on the signed-in account.
|
|
209
|
-
|
|
210
|
-
Use `--resume <root-thread-id>` to continue a saved Mission root. Without it, a single unfinished Codex
|
|
211
|
-
Mission in this repository is selected automatically; multiple unfinished roots require an explicit
|
|
212
|
-
selection. Resume restores native messages and tool evidence, including child sessions. Completed native
|
|
213
|
-
receipts feed the existing reservation reconciliation without replaying tool execution. A new Task stopped
|
|
214
|
-
before child binding can be recovered when the exact terminal parent thread/turn/call/input matches and
|
|
215
|
-
the prior adapter is closed or its Linux process identity proves it is gone. The host retains that proof in
|
|
216
|
-
the existing Mission and reconciles the reservation before another model prompt; repeated recovery does
|
|
217
|
-
not settle it again. Concurrent recovery claims are serialized.
|
|
218
|
-
|
|
219
|
-
A newly created ordinary Worker also saves its exact dispatch binding before the child prompt is sent.
|
|
220
|
-
If its parent Task result is lost, recovery can match that binding to the child's single, fully loaded,
|
|
221
|
-
completed native turn. The initial recovery path supports leaf Workers only: later turns, nested Tasks,
|
|
222
|
-
untracked native children, missing bindings and ambiguous identities remain unknown. Recovery restores
|
|
223
|
-
execution completion for the existing settlement path; missing validation evidence is a process defect,
|
|
224
|
-
not an invented PASS. It does not rerun the implementation. Continue with only the work or validation
|
|
225
|
-
still needed for normal Operator acceptance.
|
|
226
|
-
|
|
227
|
-
SIGINT/SIGTERM closes the adapter and exits 130/143. Bound children without an exact completed dispatch,
|
|
228
|
-
and commands sent without a saved terminal receipt, remain unknown and are not resent. App-server exit
|
|
229
|
-
does not prove an external command has stopped. Linux process-death checks do not prove external writer
|
|
230
|
-
quiescence either; other platforms require a clean adapter close for automatic prelaunch recovery.
|
|
231
|
-
Inspect the original executor before recovery; do not delete Mission state to force a retry.
|
|
232
|
-
Per-session usage is native thread cumulative usage (including earlier turns), not a Mission total;
|
|
233
|
-
monetary cost is unavailable and reported as `null`.
|
|
234
|
-
The existing Mission report uses observed per-turn token deltas and execution times. Cache and reasoning
|
|
235
|
-
subtotals are counted once. Its pre-terminal snapshot excludes final-answer generation, while final JSON
|
|
236
|
-
retains native thread totals. Terminal accounting observations survive cold reload in the existing Mission;
|
|
237
|
-
missing baselines remain unavailable. Native turn aggregates do not establish model-request counts, API
|
|
238
|
-
costs or remaining subscription allowance, so these values are not inferred.
|
|
239
|
-
CLI exit 0 means Operator acceptance; exit 1 means incomplete work or an execution error; exit 2
|
|
240
|
-
means invalid arguments. A native turn completing alone is not acceptance.
|
|
241
|
-
The CLI writes bounded progress records to stderr: tool start/completion, commands and exit codes,
|
|
242
|
-
replan reasons, next actions, child identity, and public model commentary. Final JSON remains on stdout.
|
|
243
|
-
|
|
244
|
-
For a manifest-bound mission, use the SDK `runCodexMission(...)` or the CLI:
|
|
86
|
+
Use **OpenCode V2**. In the target project:
|
|
245
87
|
|
|
246
88
|
```sh
|
|
247
|
-
|
|
89
|
+
npm install --save-dev sortie-dogs@0.13.13
|
|
90
|
+
npx --no-install sortie-dogs init .
|
|
248
91
|
```
|
|
249
92
|
|
|
250
|
-
|
|
251
|
-
protected evidence capture, settlement, and repository-wide scope leases. It does not emulate
|
|
252
|
-
OpenCode plugin hooks, install Codex assets, switch models implicitly, or copy/import/modify OpenCode
|
|
253
|
-
settings and sessions. Codex command and file events are post-execution observations, not a claimed
|
|
254
|
-
pre-execution write guard. The stdio transport is the supported initial integration; experimental
|
|
255
|
-
WebSocket transport is intentionally out of scope.
|
|
256
|
-
|
|
257
|
-
The manifest-run CLI handles SIGINT/SIGTERM by requesting interruption and cleaning up its
|
|
258
|
-
app-server (exit 130/143). SDK callers can pass an `AbortSignal` as `signal`. Before
|
|
259
|
-
turn dispatch, cancellation settles the reservation and releases its lease. After
|
|
260
|
-
dispatch, app-server exit or an interrupt acknowledgement does not prove that an
|
|
261
|
-
externally hosted command stopped. Cancellation or transport loss therefore preserves
|
|
262
|
-
the active goal/reservation as an unknown outcome and stops lease heartbeats without
|
|
263
|
-
claiming completion. SIGKILL can leave the same unresolved state.
|
|
93
|
+
Completely restart OpenCode, then run:
|
|
264
94
|
|
|
265
|
-
|
|
266
|
-
|
|
267
|
-
|
|
268
|
-
delete the ledger or automatically resend the prompt. This is a Codex admission
|
|
269
|
-
check; it does not establish quiescence of external command executors.
|
|
95
|
+
```text
|
|
96
|
+
/sortie-v010 Implement and verify the requested change
|
|
97
|
+
```
|
|
270
98
|
|
|
271
|
-
|
|
99
|
+
Selecting `dog-operator` starts the same workflow. `init` registers the V2 plugin, installs Mission
|
|
100
|
+
assets and sets subagent depth to at least two while preserving unrelated settings. Existing
|
|
101
|
+
`sortie-dogs/server` bridges are reused. Internal `dogs-coordinator` / `*-v010` roles are not entrypoints.
|
|
272
102
|
|
|
273
|
-
|
|
274
|
-
|
|
103
|
+
The `v010` command, profile and configuration names remain for compatibility; they do **not** mean
|
|
104
|
+
v0.10 is installed. [OpenCode setup and configuration](docs/configuration.md) ·
|
|
105
|
+
[日本語](docs/guide-ja.md) · [简体中文](docs/guide-zh-CN.md).
|
|
275
106
|
|
|
276
|
-
|
|
107
|
+
## Mission workflow
|
|
277
108
|
|
|
278
|
-
```
|
|
279
|
-
|
|
280
|
-
|
|
109
|
+
```text
|
|
110
|
+
Known unit: User → Operator → Worker → Validation → Review / Correction → Acceptance
|
|
111
|
+
Discovery: User → Operator → Coordinator → Worker → Validation → Review / Correction → Acceptance
|
|
281
112
|
```
|
|
282
113
|
|
|
283
|
-
|
|
284
|
-
|
|
114
|
+
1. **Operator** saves original requirements and owns user decisions and final acceptance.
|
|
115
|
+
2. **Worker** investigates, implements, runs meaningful checks and performs requested Git delivery
|
|
116
|
+
in the same useful unit. A known unit/check can go directly to Worker without a planning handoff.
|
|
117
|
+
3. **Coordinator**, when needed, investigates or decomposes work, reconciles in-request scope and
|
|
118
|
+
dispatches units. It can also implement and validate directly.
|
|
119
|
+
4. **Review** is initially independent for high-risk changes; low-risk skips are explicit.
|
|
120
|
+
Findings can be corrected and validated by the same Reviewer, followed by recorded author
|
|
121
|
+
self-recheck. That is not a second independent PASS.
|
|
122
|
+
5. **Acceptance** compares the original request with actual source, checks and review evidence.
|
|
123
|
+
Only the succeeded receipt authorizes DONE and the measured return report.
|
|
285
124
|
|
|
286
|
-
|
|
287
|
-
|
|
288
|
-
"plugins": ["sortie-dogs"],
|
|
289
|
-
"experimental": { "subagent_depth": 2 }
|
|
290
|
-
}
|
|
291
|
-
```
|
|
125
|
+
The default Mission profile is serial. Background responsiveness and status observations do not
|
|
126
|
+
add parallel writers, replay unknown commands or imply completion. More agents are not the goal.
|
|
292
127
|
|
|
293
|
-
|
|
294
|
-
|
|
128
|
+
For **run-once / result-collection** requests, a terminal failure can be the requested result:
|
|
129
|
+
validate and report it without an unauthorized repair/rerun loop. Acceptance preserves the real
|
|
130
|
+
operation exit and `execution-failed` outcome; it never turns a failed benchmark into a successful
|
|
131
|
+
one. Required success checks, missing commands and active/unstarted operations still block completion.
|
|
295
132
|
|
|
296
|
-
|
|
133
|
+
[Mission tools and validation policy](docs/configuration.md#mission-tools) ·
|
|
134
|
+
[Operation efficiency](docs/mission-operation-efficiency.md) ·
|
|
135
|
+
[Review policy](docs/quality-first-review.md)
|
|
297
136
|
|
|
298
|
-
|
|
299
|
-
/sortie-v010 <task>
|
|
300
|
-
```
|
|
137
|
+
## Default models
|
|
301
138
|
|
|
302
|
-
|
|
303
|
-
|
|
304
|
-
|
|
305
|
-
|
|
306
|
-
`
|
|
307
|
-
existing local bridge loads enforcement and model routing. OpenCode can reload watched configuration,
|
|
308
|
-
but replacing an installed dependency may require a full restart. A new chat session alone does not
|
|
309
|
-
prove the newly installed plugin is loaded.
|
|
310
|
-
|
|
311
|
-
## v0.13.10 runtime updates
|
|
312
|
-
|
|
313
|
-
PR #169 connects natural-language Windows Codex Missions to the same multi-role lifecycle using
|
|
314
|
-
existing PowerShell 7, preserves command failure exits and native resume settings, and maps the
|
|
315
|
-
Luna-fast Worker alias to native Luna/max with the separate priority tier. Directory reads list
|
|
316
|
-
direct entries instead of requiring filename guesses. OpenCode model settings and native permission
|
|
317
|
-
profiles remain unchanged. PR #168 provides one host-aware Anko runner, reuses unchanged setup,
|
|
318
|
-
removes mandatory paid diagnostic probes, and preserves prior attempts, unknown usage and service
|
|
319
|
-
evidence. Its recorded v0.13.9 Anko completion and PR #169's Windows Codex host-executor completion
|
|
320
|
-
are historical fixed-candidate observations, not new v0.13.10 benchmark or native sandbox results.
|
|
321
|
-
See [release notes](docs/release-v0.13.10.md).
|
|
322
|
-
|
|
323
|
-
## v0.13.9 runtime updates (retained)
|
|
324
|
-
|
|
325
|
-
PRs #163–#166 connect native Codex sessions to the existing Mission lifecycle, add the explicit
|
|
326
|
-
`$sortie-dogs` skill and `codex init`, restore operation Reviewer correction within existing write
|
|
327
|
-
scope, and reconcile native terminal observations without reviving completed work from delayed events.
|
|
328
|
-
Codex settings and authentication remain separate from OpenCode. Real Ubuntu Codex acceptance
|
|
329
|
-
evidence and host-specific limits remain in [the acceptance record](docs/codex-mission-acceptance-20261006.md);
|
|
330
|
-
they are not a new all-host or performance benchmark. Release CLI proof observes actual OpenCode
|
|
331
|
-
Worker startup and model identity only, not Mission completion. See [release notes](docs/release-v0.13.9.md).
|
|
332
|
-
|
|
333
|
-
## v0.13.8 runtime updates
|
|
334
|
-
|
|
335
|
-
PR #161 captures native validation bindings through the readable resolved target of a Windows
|
|
336
|
-
directory junction instead of enumerating its failing logical alias. Logical evidence labels,
|
|
337
|
-
link identity and complete target bytes remain bound; unchanged readable junctions retain their
|
|
338
|
-
evidence identity. External-link/cycle handling, actual source/output freshness and Review remain
|
|
339
|
-
unchanged. No new approval, permissions or model changes. The PR's isolated native Windows fixture
|
|
340
|
-
reached independent Review PASS and accepted completion after resuming the same Mission, without
|
|
341
|
-
rerunning its successful check; this does not claim original product recovery or uninterrupted
|
|
342
|
-
single-turn completion. Release receipts: `_testenv/releases/0.13.8/`;
|
|
343
|
-
see [release notes](docs/release-v0.13.8.md).
|
|
344
|
-
|
|
345
|
-
## v0.13.7 runtime updates (retained)
|
|
346
|
-
|
|
347
|
-
PRs #158/#159 prevent late write-only reports from falsely invalidating bound non-generating
|
|
348
|
-
native checks, while preserving actual source/output freshness, failed checks and legacy recipes.
|
|
349
|
-
Whole-root grants without a concrete inventory keep full-candidate freshness. Non-Git Mission roots
|
|
350
|
-
can prepare independent Review from declared nested repositories/worktrees/plain outputs, including
|
|
351
|
-
committed content; corrupt Git/HEAD errors remain visible. Operator performs authorized local recovery
|
|
352
|
-
in the same turn without another approval, changed requirements or reset spending. Ubuntu's Git
|
|
353
|
-
filesystem-boundary diagnostic is recognized alongside the standard non-repository error.
|
|
354
|
-
The PR native observation reached Reviewer startup, not final acceptance. Release receipts:
|
|
355
|
-
`_testenv/releases/0.13.7/`; see [release notes](docs/release-v0.13.7.md).
|
|
356
|
-
|
|
357
|
-
## v0.13.6 runtime updates (retained)
|
|
358
|
-
|
|
359
|
-
PR #156 references recorded formal PASS evidence in acceptance summaries instead of fetching
|
|
360
|
-
the same successful unit's native history again for display. Missing/failed proof retains history
|
|
361
|
-
diagnostics; historical evidence is not current freshness or acceptance. Identical external artifact
|
|
362
|
-
inventories are shared only inside one snapshot refresh, never across later calls. Operator guidance
|
|
363
|
-
proceeds to acceptance when current evidence covers the request and no concrete gap remains.
|
|
364
|
-
Existing freshness/review guards, provider/model selection and cache-prefix machinery remain unchanged.
|
|
365
|
-
No end-to-end speedup or live token/cache-hit improvement was measured. Release receipts:
|
|
366
|
-
`_testenv/releases/0.13.6/`; see [release notes](docs/release-v0.13.6.md).
|
|
367
|
-
|
|
368
|
-
## v0.13.5 runtime updates (retained)
|
|
369
|
-
|
|
370
|
-
PR #154 preserves stable model instructions and appends changed host state at native history
|
|
371
|
-
boundaries, with complete current-state reconstruction after compaction. Genuine Reviewer tools
|
|
372
|
-
have deterministic read-only-first ordering; the real correction-permission transition still remains.
|
|
373
|
-
The final candidate restores v3 behavioral review and combined assignment/findings while retaining
|
|
374
|
-
exact instruction discovery, inherited Reviewer formal-command delivery, literal local shell-file
|
|
375
|
-
scope reconciliation and explicit repository-root read/write scope support. New whole-project
|
|
376
|
-
captures avoid bookkeeping-only invalidation; legacy evidence and real source/artifact freshness remain.
|
|
377
|
-
|
|
378
|
-
The final pre-release cycle 11 Anko sample passed official local score 1 and selected public probes 9/9
|
|
379
|
-
in 26m1.866s at estimated $1.41968316. It was faster but 15.762% costlier than the earlier v3 sample,
|
|
380
|
-
not combined cost-preserving optimization or a general quality/speed guarantee. Release receipts:
|
|
381
|
-
`_testenv/releases/0.13.5/`. See [release notes](docs/release-v0.13.5.md) and
|
|
382
|
-
[candidate tradeoffs/failures](docs/cache-prefix-loop-20261003.md). Native Worker startup is not completion.
|
|
383
|
-
|
|
384
|
-
## v0.13.4 runtime updates (retained)
|
|
385
|
-
|
|
386
|
-
PR #152 restores same-session Coordinator/Operator implementation and formal validation through
|
|
387
|
-
`plan_units(executor="self")`, `start_direct_unit` and `finish_direct_unit`. Known single-unit work can
|
|
388
|
-
combine Mission start and planning; Luna Fast/max Worker routing remains the default. Full generated
|
|
389
|
-
contracts remain visible when an explicit Read line range covers the file. Native background
|
|
390
|
-
responsiveness remains: a launch acknowledgement or idle root is not Mission completion.
|
|
391
|
-
|
|
392
|
-
The initial independent Reviewer can investigate, correct, formally validate, deliver and self-recheck
|
|
393
|
-
continuously in its original native Task. Later Major/Medium findings accumulate in that same correction
|
|
394
|
-
context. Author self-recheck remains `self-rechecked`, `independent=false`, never independent `PASS`.
|
|
395
|
-
A different Reviewer is conditional on concrete residual Major risk; unresolved Major/Medium findings
|
|
396
|
-
still block acceptance. Operator owns final comparison and receipt.
|
|
397
|
-
|
|
398
|
-
Inherited compiler scratch no longer falsely invalidates broad-scope formal proof. Host-observed Git
|
|
399
|
-
delivery, caller-setting review and test-composition guidance reduce avoidable detours. Saved host
|
|
400
|
-
completion cards remain in tool history/UI, while outgoing V2 model/compaction context omits only their
|
|
401
|
-
presentation body, retaining receipt and evidence identities.
|
|
402
|
-
|
|
403
|
-
Two runs of the same fixed pre-release v8 package completed the original Anko task with official local
|
|
404
|
-
score 1 (F2P 9/9, P2P 94/94) and selected public probes 9/9. Times were 26m49s and 25m55s, costs
|
|
405
|
-
$1.47908048 and $1.54610816; each saved more than seven minutes versus the recorded v5 sample.
|
|
406
|
-
These limited same-task observations do not establish general speedup or a new SWE-bench score.
|
|
407
|
-
Release preflight, full tests, fixed-commit Windows CI and native Worker-start receipts are retained in
|
|
408
|
-
`_testenv/releases/0.13.4/`; startup/model identity is not task completion. See the
|
|
409
|
-
[release notes](docs/release-v0.13.4.md) and [quality-loop evidence](docs/nightly-quality-loop-20261002.md).
|
|
139
|
+
- **Operator / Coordinator / Reviewer / Advisor:** GPT-6.1 Sol, `xhigh`.
|
|
140
|
+
- **Worker / Scout:** Luna Fast, `max`.
|
|
141
|
+
- **OpenCode:** `openai/gpt-6.1-sol#xhigh` and `openai/gpt-6-luna-fast#max`.
|
|
142
|
+
- **Codex:** native `gpt-6.1-sol` / `xhigh`; the Luna-fast alias maps to native
|
|
143
|
+
`gpt-6-luna` / `max` with separate `serviceTier: "priority"` (Fast).
|
|
410
144
|
|
|
411
|
-
|
|
145
|
+
Explicit host model selections remain authoritative. OpenCode routing and Codex native settings
|
|
146
|
+
are separate; unsupported native models fail rather than silently changing route. Codex Fast uses
|
|
147
|
+
subscription limits faster than Standard. Subscription USD cost and remaining allowance are not
|
|
148
|
+
inferred from token totals.
|
|
412
149
|
|
|
413
|
-
|
|
414
|
-
use Operator → Coordinator → Worker for actual discovery or decomposition:
|
|
415
|
-
|
|
416
|
-
- `dog-operator` states a few requirements/negative constraints and owns user decisions and final acceptance.
|
|
417
|
-
The host saves the original user message verbatim.
|
|
418
|
-
- Hidden `dogs-coordinator` owns investigation, unit declarations, Worker/Scout/Advisor/Reviewer dispatch,
|
|
419
|
-
in-request write-scope extensions, and corrections. It can read/search and run confirmation shell commands;
|
|
420
|
-
it can implement and formally validate directly in its own session, or delegate a unit to Worker.
|
|
421
|
-
- `dog-worker-v010` implements a host-generated unit within its file/directory write scopes.
|
|
422
|
-
Investigation commands need no pre-registration; formal checks retain real host-recorded results.
|
|
423
|
-
- High-risk changes require an initial independent Reviewer with read/search access. Low-risk skips
|
|
424
|
-
are explicit and recorded. Reviewer-owned corrections follow the self-recheck policy above.
|
|
425
|
-
- Fast-lane can include high-risk single-unit work; it never implies a review skip.
|
|
426
|
-
- Investigation, edits, formal checks and requested Git delivery stay in the same implementing child.
|
|
427
|
-
Explicit user ordering is retained; no routine plan-approval or commit-only handoff is needed.
|
|
428
|
-
- Unit progress appears on the running Task without stopping Coordinator or prompting Operator.
|
|
429
|
-
|
|
430
|
-
The `v010` Mission profile is serial by design; background responsiveness does not add parallel writers.
|
|
431
|
-
The stable profile's Luna fabric and parallel integration path are not exposed in this profile. More agents are not a
|
|
432
|
-
goal; preserving quality while reducing unnecessary expensive work is.
|
|
150
|
+
## Measured results
|
|
433
151
|
|
|
434
|
-
###
|
|
152
|
+
### Codex end-to-end completion
|
|
435
153
|
|
|
436
|
-
|
|
437
|
-
|
|
438
|
-
| Candidate / benchmark adapter | Official resolved | Empty patches | Run conditions | Details |
|
|
439
|
-
| --- | ---: | ---: | --- | --- |
|
|
440
|
-
| v0.10.6 candidate (`859c396`) | 4 / 23 (17.4%) | 5 | Four inference slots | [Handoff](docs/swebench-handoff-2026-09-21.md) |
|
|
441
|
-
| Frozen v0.10.14 build | 6 / 23 (26.1%) | — | Four slots; $1.50/task | [Per-task results](docs/benchmark-v0.10.14-dev23.md) |
|
|
442
|
-
| v0.12.8 (`e0f8cef` adapter) | 5 / 23 (21.7%) | 9 | Four slots; budget amended across two batches; 30-minute timeout | [Campaign](docs/swebench-v0128-dev23-2026-09-26.md) |
|
|
443
|
-
| v0.12.8 (`84ccdf1` main adapter) | 6 / 23 (26.1%) | 2 | Fresh 23-task run; four slots; 30-minute timeout | [Main-integrated run](docs/swebench-main-84ccdf1-dev23-2026-09-26.md) |
|
|
444
|
-
| v0.12.15 (`0c9690d`) | 5 / 23 (21.7%) | 3 | Fresh 23-task ext4 retry; four slots; 40-minute timeout | [Campaign](docs/swebench-v01215-dev23-2026-09-27.md) |
|
|
445
|
-
| v0.12.16 (`9b05a34` release; `b1a6c0e` runner) | 4 / 23 (17.4%) | 5 | Fresh 23-task run; eight slots; 40-minute timeout | [Campaign](docs/swebench-v01216-dev23-2026-09-27.md) |
|
|
446
|
-
| v0.12.19 (`24f5386` release; matched rerun) | 8 / 23 (34.8%) | 0 | Eight slots; effective $2/instance; 40-minute timeout; one inference timeout | [Official result and provenance](docs/benchmarks/swebench-v01220-operation-observability-2026-09-28.md) |
|
|
447
|
-
| v0.12.20 (`628eb81` release) | 7 / 23 (30.4%) | 0 | Eight slots; effective $2/instance; 40-minute timeout | [Official result and caveat](docs/benchmarks/swebench-v01220-operation-observability-2026-09-28.md) |
|
|
448
|
-
| v0.12.25 (`49eb1e4` release) | 7 / 23 (30.4%) | 0 | Eight slots; $2/instance; $30 total cap; 20-minute progress check / 40-minute hard maximum | [Comparison baseline](#v0131-dev23-2026-09-30) |
|
|
449
|
-
| v0.13.1 (`d19e8be` release; 2026-09-30) | **8 / 23 (34.8%)** | 0 | Eight slots; $2/instance; $46 total cap; 20-minute progress check / 40-minute hard maximum; GPT-6.1 Sol + Luna Fast | [Run summary](#v0131-dev23-2026-09-30) |
|
|
450
|
-
|
|
451
|
-
Every row has 23 submitted official predictions; an empty patch counts against
|
|
452
|
-
the score, not as a missing evaluation. The v0.10.14 report does not separately
|
|
453
|
-
summarize empty patches. The v0.12.15 row is the separately approved fresh
|
|
454
|
-
run after an initial `/tmp` quota failure, not an additional score for that
|
|
455
|
-
failed attempt. Inference completion is **not** official resolution.
|
|
456
|
-
Budgets, runtime/adapter versions, and execution conditions changed between
|
|
457
|
-
campaigns, so this table is a history of observed results, not a controlled
|
|
458
|
-
head-to-head comparison or a general success-rate claim. The v0.10.14 run
|
|
459
|
-
estimated $15.75 in model cost and a 15.2-minute median agent runtime.
|
|
460
|
-
The v0.12.16 run is the first eight-slot inference score in this table;
|
|
461
|
-
five runners timed out, and this does not establish a model-quality regression
|
|
462
|
-
against runs with different concurrency and conditions.
|
|
463
|
-
The v0.12.19 row is the corrected run with an effective $2 per-instance cap;
|
|
464
|
-
an earlier v0.12.19 run scored 7/23 but had no effective per-instance cap and
|
|
465
|
-
is not a same-condition comparison. The v0.12.20 run lost the
|
|
466
|
-
`sqlfluff__sqlfluff-2419` resolution relative to the corrected run; this
|
|
467
|
-
single run-to-run difference does not establish causation.
|
|
468
|
-
|
|
469
|
-
Historical qualification references remain in [benchmark reference](docs/benchmark-reference.md).
|
|
154
|
+
**Codex integration has completed real Missions on Ubuntu and Windows.** The Windows large-write
|
|
155
|
+
Anko trial reached final Operator acceptance after Worker validation and same-Reviewer correction:
|
|
470
156
|
|
|
471
|
-
|
|
157
|
+
- SDK exit **0**, `accepted: true`, Mission **completed**.
|
|
158
|
+
- **190/190** terminal command receipts; no unresolved command. Large generated-parser writes completed.
|
|
159
|
+
- Separate local replay of official tests: reward **1**, F2P **9/9**, P2P **94/94**.
|
|
160
|
+
- Elapsed **30m08.739s**, including cleanup; ChatGPT subscription monetary cost unavailable.
|
|
472
161
|
|
|
473
|
-
|
|
474
|
-
|
|
475
|
-
|
|
476
|
-
|
|
477
|
-
|
|
478
|
-
|
|
479
|
-
- The scored row is the user-requested fresh run after a host restart. The
|
|
480
|
-
interrupted initial run is excluded from this score; the fresh run made one
|
|
481
|
-
attempt per instance with no inference retry.
|
|
482
|
-
- The dataset revision (`6ec7bb89b9342f664a54a6e0a6ea6501d3437cc2`), public rows,
|
|
483
|
-
and all 23 official evaluation image IDs match the v0.12.25 run. Both used
|
|
484
|
-
`official-image-testbed` and sequential official scoring.
|
|
485
|
-
- Operator/Coordinator/Reviewer/Advisor defaults changed to
|
|
486
|
-
`openai/gpt-6.1-sol#xhigh`. Actual task Workers remained
|
|
487
|
-
`openai/gpt-6-luna-fast#max`, observed across all 23 instances. The harness and
|
|
488
|
-
total budget also changed, so the extra resolution cannot be attributed to
|
|
489
|
-
the model change alone.
|
|
490
|
-
- Inference ended with 20 normal completions, two timeouts
|
|
491
|
-
(`pvlib__pvlib-python-1154`, `sqlfluff__sqlfluff-1763`) and one agent failure
|
|
492
|
-
(`pvlib__pvlib-python-1854`). All patches, including stopped attempts, were
|
|
493
|
-
officially scored: 23 completed evaluations, zero empty patches and zero
|
|
494
|
-
official evaluation errors or infrastructure failures.
|
|
495
|
-
- Known estimated inference cost: **$17.73**; separate unknown-usage hold:
|
|
496
|
-
**$1.98**, not counted as known expense. Inference wall time was about
|
|
497
|
-
**81 minutes**, followed by **7.2 minutes** of official scoring.
|
|
498
|
-
- Fixed release commit: `d19e8be0d21180cc23ad2ae4b853d846a18e77bc`;
|
|
499
|
-
package SHA-256: `99300ceec0c3eee4fa1f984fed50d15d58b2ff4455df0041850ecd63514b0a51`.
|
|
500
|
-
OpenCode **2.0.20**, official harness **5.0.2**. Local evidence is retained in
|
|
501
|
-
`_testenv/swebench-v0131-dev23-20260930-r2/result-summary.json`; generated
|
|
502
|
-
predictions, databases and raw logs are not committed.
|
|
503
|
-
|
|
504
|
-
## Mission tools
|
|
505
|
-
|
|
506
|
-
1. `start_mission`: Operator supplies concise requirements; the host saves original messages. A known single unit
|
|
507
|
-
can include `unit` to combine start/planning and return its configured Worker task.
|
|
508
|
-
2. `plan_units`: Operator or Coordinator supplies title, objective, file/directory scopes and formal checks. The host generates
|
|
509
|
-
IDs, handoff, manifest, proof mapping and the ready Worker task. `executor="self"` keeps execution in the
|
|
510
|
-
same controller session; `start_direct_unit` / `finish_direct_unit` retain observed formal-check freshness.
|
|
511
|
-
3. `operator_next`: advance serial units. `expand_unit` reconciles required in-request outputs while
|
|
512
|
-
preserving the same Task. A reasoned `plan_units` correction or `retry_mission_unit` handles ordinary
|
|
513
|
-
unit recovery under the original requirements and cumulative budget.
|
|
514
|
-
4. `review_mission`: generate the independent review packet from source, requirements and observed checks;
|
|
515
|
-
dispatch its Reviewer task for high-risk changes or record a low-risk skip.
|
|
516
|
-
5. `repair_review`: record findings and continue correction in the running original Reviewer Task, or resume
|
|
517
|
-
that same native session. Later findings accumulate; run inherited checks/requested Git delivery and
|
|
518
|
-
explicitly report `SELF_RECHECKED` in that same Task. Legacy
|
|
519
|
-
`CORRECTION_READY` alone requires a same-author read-only fallback through `review_mission`.
|
|
520
|
-
6. `submit_mission`: Coordinator returns a completion candidate, user-only decision, or proven external/scope/budget blocker.
|
|
521
|
-
7. `complete_mission`: Operator compares the original request, source and evidence, then explicitly accepts.
|
|
522
|
-
Only a succeeded receipt authorizes DONE and the measured 🐾 return report.
|
|
523
|
-
|
|
524
|
-
All tool names use the `sortie_v010_` prefix. Prior proposal/plan-repair tools remain in the compatibility
|
|
525
|
-
implementation but are hidden from the normal Mission tool list. Operator/Coordinator handle in-request
|
|
526
|
-
path reconciliation without a new user approval. Changes beyond the original requirements or cumulative
|
|
527
|
-
budget return to Operator/user. `EVIDENCE_GAPS` is an advisory limitation, not Review `PASS` or an
|
|
528
|
-
automatic extra review; failed or missing required checks still prevent acceptance.
|
|
529
|
-
|
|
530
|
-
Durable profile state and hash-bound task references support restart and
|
|
531
|
-
compaction recovery without reconstructing criteria from summary prose. Stale,
|
|
532
|
-
foreign-root, or changed references are rejected. Requested `git add <paths>` and `git commit -m ...`
|
|
533
|
-
use the actual source write scope, not a fabricated `.git/**` scope. An optional host-managed Git
|
|
534
|
-
lifecycle also retains its branch, commit and post-commit boundaries. Neither mode grants arbitrary
|
|
535
|
-
Git, force push, release or publication authority.
|
|
536
|
-
|
|
537
|
-
### Progress and acceptance evidence
|
|
538
|
-
|
|
539
|
-
`sortie_v010_operator_status` keeps original requests, formal command/exit/timing observations,
|
|
540
|
-
review disposition and recorded delivery in a compact Mission view. `{ "view": "progress" }` exposes
|
|
541
|
-
the current unit, completed/total units, host budget and next action; `{ "view": "full" }` or
|
|
542
|
-
`details_ref` provides full snapshot diagnostics. Unknown clean state or a failed commit is not delivery
|
|
543
|
-
success. Reading progress does not dispatch, retry or accept work; use native completion notifications
|
|
544
|
-
instead of polling. Existing status reconciliation can recover a missed child terminal event.
|
|
162
|
+
The run used native Codex models/threads with an **explicitly authorized SDK parent-host executor**.
|
|
163
|
+
It does not claim native Windows sandbox repair, Docker-equivalent scoring or a hosted leaderboard
|
|
164
|
+
submission. Its fixed development package preceded v0.13.11, which includes the changes; the
|
|
165
|
+
published release was not rerun or assigned that score. These are completion observations, not a
|
|
166
|
+
general speed or success-rate guarantee.
|
|
545
167
|
|
|
546
|
-
|
|
168
|
+
[Windows run and scoring](docs/codex-write-status-anko-20261007.md) ·
|
|
169
|
+
[Windows models/shell evidence](docs/codex-windows-luna-fast-20261007.md) ·
|
|
170
|
+
[Ubuntu acceptance](docs/codex-mission-acceptance-20261006.md)
|
|
547
171
|
|
|
548
|
-
###
|
|
549
|
-
|
|
550
|
-
The default package entry is the `v010` Mission profile:
|
|
551
|
-
|
|
552
|
-
- Command: `/sortie-v010`
|
|
553
|
-
- Primary agent: `dog-operator`
|
|
554
|
-
- Project settings: `.opencode/sortie-dogs-v010.json`
|
|
555
|
-
- Global settings: `~/.config/opencode/sortie-dogs-v010.json`
|
|
556
|
-
- JSON environment override: `SORTIE_DOGS_V010_CONFIG`
|
|
557
|
-
- Runtime state: `.sortie-dogs-v010/`
|
|
558
|
-
- Installed asset marker: `.opencode/sortie-dogs-v010.version`
|
|
559
|
-
|
|
560
|
-
For a global install, the marker is `<OpenCode config root>/sortie-dogs-v010.version`.
|
|
561
|
-
|
|
562
|
-
Precedence is built-in defaults, global file, project file, environment JSON,
|
|
563
|
-
then plugin factory options. Unknown properties or invalid types are rejected.
|
|
564
|
-
Use external `v010` role names such as `dog-operator`, `dogs-coordinator`, and
|
|
565
|
-
`dog-reviewer-v010` in `modelRouting`; do not also declare their stable aliases.
|
|
566
|
-
|
|
567
|
-
Example `.opencode/sortie-dogs-v010.json`:
|
|
568
|
-
|
|
569
|
-
```json
|
|
570
|
-
{
|
|
571
|
-
"validationProfile": "balanced",
|
|
572
|
-
"readOnlyTools": ["my_mcp_search"],
|
|
573
|
-
"freeTierFallbackModels": ["opencode/deepseek-v4-flash-free"],
|
|
574
|
-
"modelRouting": {
|
|
575
|
-
"dog-operator": {
|
|
576
|
-
"preferred": { "model": "provider/model", "variant": "high" }
|
|
577
|
-
},
|
|
578
|
-
"dogs-coordinator": {
|
|
579
|
-
"preferred": { "model": "provider/model", "variant": "deep" }
|
|
580
|
-
}
|
|
581
|
-
},
|
|
582
|
-
"modelCatalog": {
|
|
583
|
-
"project": [
|
|
584
|
-
{ "model": "provider/model", "variants": ["high", "deep"] }
|
|
585
|
-
]
|
|
586
|
-
},
|
|
587
|
-
"continuation": {
|
|
588
|
-
"enabled": true,
|
|
589
|
-
"taskWatchdogMilliseconds": 300000
|
|
590
|
-
}
|
|
591
|
-
}
|
|
592
|
-
```
|
|
172
|
+
### SWE-bench evaluation
|
|
593
173
|
|
|
594
|
-
|
|
595
|
-
|
|
596
|
-
|
|
597
|
-
|
|
598
|
-
|
|
599
|
-
- `readOnlyTools`: additional host-specific tools known not to mutate project
|
|
600
|
-
files. Values accumulate across configuration layers. Unknown tools are denied
|
|
601
|
-
in a bound worker session.
|
|
602
|
-
- `modelRouting`: preferred and ordered fallback targets by external profile role.
|
|
603
|
-
- `modelCatalog`: available `project` and `global` model/variant declarations.
|
|
604
|
-
- `freeTierFallbackModels`: ordered global last-resort model IDs. Default:
|
|
605
|
-
`opencode/deepseek-v4-flash-free`; `[]` disables this fallback.
|
|
606
|
-
- `dedicatedWorkerModel`: canonical stable serial target, default
|
|
607
|
-
`openai/gpt-6.1-sol` / `medium`. The Mission profile supplies its explicit
|
|
608
|
-
role routes below; do not infer its Worker route from this stable setting.
|
|
609
|
-
- `consultation.strategy`: fixed advisor identity, optional `required`, and
|
|
610
|
-
positive `maxCallsPerCandidate`; default one call and not required.
|
|
611
|
-
- `consultation.sourceReview`: risk-based review with `maxCallsPerCandidate`
|
|
612
|
-
default `1` and `maxArtifactBytes` default/maximum `30720`. Unavailable review
|
|
613
|
-
blocks only when review is required.
|
|
614
|
-
- `continuation.enabled`: default `true`.
|
|
615
|
-
- Automatic continuation has no turn-count ceiling. The legacy positive-integer
|
|
616
|
-
`continuation.maxAutoContinues` setting is accepted but ignored.
|
|
617
|
-
- `continuation.taskWatchdogMilliseconds`: root inactivity while an implementation
|
|
618
|
-
Task is outstanding; default `300000`, valid range `10..1800000`.
|
|
619
|
-
- `continuation.summarizeModel`: optional explicit compaction model; omission
|
|
620
|
-
reuses the latest observed root model.
|
|
621
|
-
- `validationProfile`: `fast`, `balanced`, or `assurance`; default `balanced`.
|
|
622
|
-
- `reflection`: enabled by default for `run`, `project` and `global` layers, with at most three
|
|
623
|
-
entries / 500 estimated tokens injected. Root Operator can use `sortie_v010_reflection` to retain
|
|
624
|
-
verified process causes/preventions for later turns and sessions; this is not model training.
|
|
625
|
-
Storage and managed blocks are separate from stable. Set `reflection.enabled` to `false` to disable.
|
|
626
|
-
|
|
627
|
-
The Mission host owns handoff and manifest controls under
|
|
628
|
-
`.sortie-dogs-v010/contracts/`. Do not create a legacy root
|
|
629
|
-
`operation-manifest.json` for this profile and do not edit generated controls.
|
|
630
|
-
Delete `.sortie-dogs-v010/` only when no Sortie run is active.
|
|
631
|
-
|
|
632
|
-
### Validation policy
|
|
633
|
-
|
|
634
|
-
`validationProfile` chooses supplementary non-canonical depth:
|
|
635
|
-
|
|
636
|
-
- `fast`: static checks
|
|
637
|
-
- `balanced`: targeted checks
|
|
638
|
-
- `assurance`: related checks
|
|
639
|
-
|
|
640
|
-
It does not replace meaningful declared formal checks or user/project-required broad validation.
|
|
641
|
-
Batch related edits, run focused checks, then execute every declared formal check in order on the
|
|
642
|
-
stable candidate. The implementing Worker or admitted correcting Reviewer runs those commands;
|
|
643
|
-
canonical/full-suite `owner=coordinator` is evidence accounting, not a requirement for root execution.
|
|
644
|
-
|
|
645
|
-
Keep required broad checks for the final integrated candidate. Reuse valid unchanged evidence only
|
|
646
|
-
when the contract permits; identity includes candidate, command, environment, scope and owner.
|
|
647
|
-
Required repeated occurrences retain their own execution identities and cannot be skipped as duplicates.
|
|
648
|
-
Native commands, actual working directory, exit, duration and saved source bindings establish freshness.
|
|
649
|
-
Diagnostics do not substitute for formal proof. Repeat checks when changes, failures or freshness require
|
|
650
|
-
it, rather than solely because a Worker changed or documentation was edited.
|
|
651
|
-
|
|
652
|
-
### Default routes
|
|
653
|
-
|
|
654
|
-
- `dog-operator`: `openai/gpt-6.1-sol` / `xhigh`
|
|
655
|
-
- `dogs-coordinator`: `openai/gpt-6.1-sol` / `xhigh`
|
|
656
|
-
- `dog-worker-v010`: `openai/gpt-6-luna-fast` / `max`
|
|
657
|
-
- `dog-scout-v010`: `openai/gpt-6-luna-fast` / `max`
|
|
658
|
-
- `dog-reviewer-v010`: `openai/gpt-6.1-sol` / `xhigh`
|
|
659
|
-
- `dog-advisor-v010`: `openai/gpt-6.1-sol` / `xhigh`
|
|
660
|
-
|
|
661
|
-
An explicit model and variant selected in OpenCode remains authoritative for that
|
|
662
|
-
session. Child role defaults fill absent native settings and may be overridden by
|
|
663
|
-
valid profile routing. Review never silently inherits the implementation model.
|
|
664
|
-
|
|
665
|
-
## Stable compatibility profile
|
|
666
|
-
|
|
667
|
-
The earlier parallel-capable runtime remains available explicitly:
|
|
174
|
+
**SWE-bench Lite test300: 170/300 (56.67%)**, fixed **v0.12.24**, one pass@1 campaign.
|
|
175
|
+
Nine empty patches, zero official evaluation errors, frozen predictions and trajectories for all
|
|
176
|
+
300 instances. Workers ran Luna Fast/max; management/review ran GPT-6 Sol/xhigh.
|
|
177
|
+
This is a system result, **not** a Luna-only comparison or Verified/full SWE-bench score.
|
|
668
178
|
|
|
669
|
-
|
|
670
|
-
|
|
671
|
-
```
|
|
179
|
+
Confirmed inference expense **$162.99**; separate unknown-pricing budget hold **$34.60** is not
|
|
180
|
+
known expense. Official local evaluation and leaderboard acceptance are separate.
|
|
672
181
|
|
|
673
|
-
|
|
182
|
+
[Technical report](docs/benchmarks/swebench-lite-v01224-test300-2026-09-29.md) ·
|
|
183
|
+
[Public predictions/logs/trajectories](https://github.com/zufall-upon/sortie-dogs-swebench-lite-20260929) ·
|
|
184
|
+
[Result history and conditions](docs/benchmarks/swebench-lite-history.md)
|
|
674
185
|
|
|
675
|
-
|
|
676
|
-
export { SortieDogsPlugin } from "sortie-dogs/plugin/stable";
|
|
677
|
-
```
|
|
186
|
+
#### v0.13.1 dev23 (2026-09-30)
|
|
678
187
|
|
|
679
|
-
|
|
680
|
-
|
|
681
|
-
register stable and `v010` from the same package installation path in one host.
|
|
188
|
+
**8/23 (34.8%)**, versus v0.12.25's 7/23; all seven prior resolutions retained, no empty patches or
|
|
189
|
+
official evaluation errors. [Fixed conditions, stopped-run accounting and costs](docs/benchmarks/swebench-lite-history.md#v0131-dev23-2026-09-30).
|
|
682
190
|
|
|
683
|
-
|
|
191
|
+
All scores belong to their fixed candidates, **not v0.13.13**. SWE-bench remains an optional
|
|
192
|
+
measurement, not a release gate. Different budgets, models and conditions are not controlled comparisons.
|
|
684
193
|
|
|
685
|
-
|
|
194
|
+
## Configuration
|
|
686
195
|
|
|
687
|
-
|
|
688
|
-
|
|
689
|
-
sortie-dogs
|
|
690
|
-
|
|
196
|
+
**OpenCode:** project settings `.opencode/sortie-dogs-v010.json`, global settings
|
|
197
|
+
`~/.config/opencode/sortie-dogs-v010.json`, JSON override `SORTIE_DOGS_V010_CONFIG`.
|
|
198
|
+
The default Mission state is `.sortie-dogs-v010/`; generated controls are host-owned.
|
|
199
|
+
[Precedence, settings, global installation and stable compatibility](docs/configuration.md).
|
|
691
200
|
|
|
692
|
-
|
|
693
|
-
|
|
201
|
+
**Codex:** use `codex mission --help` or `CodexMissionSession`. Optional `--model`, `--effort`,
|
|
202
|
+
`--permissions`, `--trusted-pwsh` and `--resume` are explicit host choices, not OpenCode settings.
|
|
203
|
+
[SDK and recovery reference](docs/codex.md).
|
|
694
204
|
|
|
695
|
-
|
|
696
|
-
can resolve a **separate dependency** under that config root. Updating npm-global alone does not update
|
|
697
|
-
it. For that layout, also install the same release at the actual config root, then rerun global init:
|
|
205
|
+
## Updates and removal
|
|
698
206
|
|
|
699
|
-
|
|
700
|
-
npm install --prefix "$HOME/.config/opencode" sortie-dogs@0.13.10
|
|
701
|
-
sortie-dogs init --global --profile v010
|
|
702
|
-
```
|
|
207
|
+
Install the desired package version and rerun **your host's** initializer:
|
|
703
208
|
|
|
704
|
-
|
|
705
|
-
|
|
706
|
-
version pins require an explicit version change. Completely restart OpenCode after updating, then
|
|
707
|
-
check the installed package, asset marker and loaded plugin version.
|
|
209
|
+
- Codex: `npx --no-install sortie-dogs codex init .`, then restart/open a new Codex chat.
|
|
210
|
+
- OpenCode: `npx --no-install sortie-dogs init .`, then completely restart OpenCode.
|
|
708
211
|
|
|
709
|
-
|
|
212
|
+
An OpenCode config-local bridge may resolve a separate package; npm-global update alone does not
|
|
213
|
+
update it. [Global/update instructions](docs/configuration.md#global-availability).
|
|
214
|
+
The current Mission asset marker is `0.13.13-mission-context-v1`; the unchanged Codex skill marker is
|
|
215
|
+
`0.13.11-codex-skill-v2`. Installed markers alone do not prove a running host reloaded the new version.
|
|
710
216
|
|
|
711
|
-
|
|
217
|
+
There is no supported uninstall command. Remove the dependency and only known Sortie-owned paths;
|
|
218
|
+
follow the [manual removal guide](docs/uninstall.md), never delete the entire `.opencode` directory.
|
|
712
219
|
|
|
713
|
-
|
|
714
|
-
npm install --save-dev sortie-dogs@latest
|
|
715
|
-
npx sortie-dogs init .
|
|
716
|
-
```
|
|
220
|
+
## Documentation
|
|
717
221
|
|
|
718
|
-
|
|
719
|
-
|
|
720
|
-
|
|
222
|
+
- **Usage:** [Codex guide](docs/codex.md) · [OpenCode configuration](docs/configuration.md) ·
|
|
223
|
+
[日本語 OpenCodeガイド](docs/guide-ja.md) · [简体中文 OpenCode指南](docs/guide-zh-CN.md).
|
|
224
|
+
- **Design:** [Worker instructions](docs/worker-instruction-design.md) ·
|
|
225
|
+
[Quality-first review](docs/quality-first-review.md) · [Operation efficiency](docs/mission-operation-efficiency.md).
|
|
226
|
+
- **Evidence:** [Codex completion](docs/codex-write-status-anko-20261007.md) ·
|
|
227
|
+
[SWE-bench history](docs/benchmarks/swebench-lite-history.md) · [Historical local case study](docs/benchmark-reference.md).
|
|
228
|
+
- **Development:** [Testing](docs/testing.md) · [CLI testing](docs/cli-testing.md) ·
|
|
229
|
+
[Windows tests](docs/windows-tests.md) · [Release routine](docs/release-batch.md).
|
|
230
|
+
- **Releases:** [v0.13.13](docs/release-v0.13.13.md) · [v0.13.12](docs/release-v0.13.12.md) · [v0.13.11](docs/release-v0.13.11.md) · [v0.13.10](docs/release-v0.13.10.md) ·
|
|
231
|
+
[v0.13.9](docs/release-v0.13.9.md) · [All GitHub releases](https://github.com/zufall-upon/Sortie-dogs/releases).
|
|
721
232
|
|
|
722
|
-
|
|
723
|
-
marker of `0.13.10-codex-windows-v1` identifies the assets; it does not prove an already-running
|
|
724
|
-
OpenCode process has reloaded the plugin.
|
|
233
|
+
## Community
|
|
725
234
|
|
|
726
|
-
|
|
727
|
-
|
|
728
|
-
|
|
729
|
-
wildcards.
|
|
235
|
+
[Contributing](CONTRIBUTING.md) · [Code of conduct](CODE_OF_CONDUCT.md) ·
|
|
236
|
+
[Security policy / private reports](SECURITY.md) · [Accessibility](ACCESSIBILITY.md) ·
|
|
237
|
+
[Report a bug or propose a feature](https://github.com/zufall-upon/Sortie-dogs/issues/new/choose).
|
|
730
238
|
|
|
731
|
-
|
|
732
|
-
validation, global application, GitHub publication, and manual npm publication.
|
|
239
|
+
Licensed under [MIT](LICENSE).
|