sortie-dogs 0.13.2 → 0.13.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +142 -48
- package/dist/asset-version.d.ts +1 -1
- package/dist/asset-version.js +1 -1
- package/dist/core/contract-limits.d.ts +6 -0
- package/dist/core/contract-limits.js +3 -0
- package/dist/core/operator-mission.d.ts +69 -4
- package/dist/core/operator-mission.js +148 -31
- package/dist/core/operator-runtime.d.ts +36 -0
- package/dist/core/operator-runtime.js +189 -18
- package/dist/core/runtime-profile.js +3 -0
- package/dist/core/validate-schema.js +0 -1
- package/dist/core/validation-budget.d.ts +5 -0
- package/dist/core/validation-budget.js +5 -2
- package/dist/plugin/gate.d.ts +7 -1
- package/dist/plugin/gate.js +131 -24
- package/dist/plugin/index.d.ts +17 -0
- package/dist/plugin/index.js +354 -59
- package/dist/plugin/mission-review.d.ts +112 -2
- package/dist/plugin/mission-review.js +225 -36
- package/dist/plugin/native-background.d.ts +40 -0
- package/dist/plugin/native-background.js +211 -0
- package/dist/plugin/native-contract-read.d.ts +9 -0
- package/dist/plugin/native-contract-read.js +98 -0
- package/dist/plugin/profiled.js +936 -182
- package/dist/plugin/protected-snapshot.d.ts +4 -0
- package/dist/plugin/protected-snapshot.js +51 -13
- package/dist/plugin/receipt-presentation.d.ts +2 -0
- package/dist/plugin/receipt-presentation.js +48 -0
- package/dist/plugin/run-metrics.d.ts +1 -1
- package/dist/plugin/run-metrics.js +3 -2
- package/dist/plugin/runtime-bridge.d.ts +67 -2
- package/dist/plugin/v2.d.ts +40 -20
- package/dist/plugin/v2.js +347 -55
- package/dist/plugin/validation-scratch.d.ts +2 -2
- package/dist/plugin/validation-scratch.js +29 -17
- package/dist/runtime-assets-v010.js +14 -6
- package/dist/runtime-mission-assets.d.ts +5 -3
- package/dist/runtime-mission-assets.js +188 -133
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -14,6 +14,9 @@ implementation, validation, review, and model routing.
|
|
|
14
14
|
continuation, remediation, and restart.
|
|
15
15
|
- **Adaptive execution**: small work stays small; additional agents and stronger
|
|
16
16
|
models are used only when task shape or risk justifies them.
|
|
17
|
+
- **Clear Worker instructions**: give Luna a concise goal, explicit constraints
|
|
18
|
+
and completion criteria; keep orchestration bookkeeping in the harness.
|
|
19
|
+
See the [instruction design principles](docs/worker-instruction-design.md).
|
|
17
20
|
- **Coexistence**: Sortie activates only when selected and preserves normal
|
|
18
21
|
OpenCode agents, settings, and user-owned files.
|
|
19
22
|
- **Cost, time, and proof**: the objective is a verified result at the lowest
|
|
@@ -32,6 +35,11 @@ implementation, validation, review, and model routing.
|
|
|
32
35
|
Guides: [日本語](docs/guide-ja.md) · [简体中文](docs/guide-zh-CN.md) ·
|
|
33
36
|
[Testing](docs/testing.md) · [CLI testing](docs/cli-testing.md)
|
|
34
37
|
|
|
38
|
+
**Current release: [v0.13.4](https://github.com/zufall-upon/Sortie-dogs/releases/tag/v0.13.4)**
|
|
39
|
+
([release notes](docs/release-v0.13.4.md)). The default Mission runtime retains the `v010`
|
|
40
|
+
profile, command and configuration names for compatibility; these names do not mean v0.10 is installed.
|
|
41
|
+
The current asset marker is `0.13.4-reviewer-continuous-v1`.
|
|
42
|
+
|
|
35
43
|
## SWE-bench Lite: 170/300 (56.67%)
|
|
36
44
|
|
|
37
45
|
The fixed **Sortie-dogs v0.12.24** harness resolved **170 of 300 SWE-bench Lite test issues** in one pass@1 campaign, with 9 empty patches and no official evaluation errors. Every instance has a frozen prediction and an inference-time trajectory. The task Workers ran `openai/gpt-6-luna-fast#max`; operator, coordinator and review roles ran `openai/gpt-6-sol#xhigh`. This is a system result, **not** a Luna-only model comparison or a Verified/full SWE-bench score.
|
|
@@ -40,22 +48,25 @@ The fixed **Sortie-dogs v0.12.24** harness resolved **170 of 300 SWE-bench Lite
|
|
|
40
48
|
|
|
41
49
|
The single official 300-instance report and frozen predictions are hash-bound in the report. Confirmed inference expense was **$162.99**; a separate **$34.60** of usage has unknown pricing and is held against the campaign cap, **not** counted as known expense. Leaderboard registration and maintainer acceptance are separate from this official local evaluation.
|
|
42
50
|
|
|
43
|
-
|
|
51
|
+
Historical scores below belong to their fixed candidates, not v0.13.4. SWE-bench is a separate,
|
|
52
|
+
optional measurement rather than a mandatory release gate.
|
|
53
|
+
|
|
54
|
+
> **Beta:** v0.13.x is still stabilizing. Runtime behavior,
|
|
44
55
|
> configuration, and generated assets may still change before 1.0.
|
|
45
56
|
|
|
46
57
|
## Quick start
|
|
47
58
|
|
|
48
|
-
Requirements: Node.js 22.6 or newer, npm, and OpenCode.
|
|
59
|
+
Requirements: Node.js 22.6 or newer, npm, and OpenCode V2.
|
|
49
60
|
|
|
50
61
|
Run these commands in the target project:
|
|
51
62
|
|
|
52
63
|
```sh
|
|
53
|
-
npm install --save-dev sortie-dogs
|
|
64
|
+
npm install --save-dev sortie-dogs@latest
|
|
54
65
|
npx sortie-dogs init .
|
|
55
66
|
```
|
|
56
67
|
|
|
57
|
-
|
|
58
|
-
|
|
68
|
+
`init` defaults to the `v010` Mission profile and registers the OpenCode V2 plugin in
|
|
69
|
+
`.opencode/opencode.json(c)`, preserving unrelated settings. It sets subagent depth to at least two:
|
|
59
70
|
|
|
60
71
|
```json
|
|
61
72
|
{
|
|
@@ -64,37 +75,72 @@ two-level subagent depth to `.opencode/opencode.json(c)`, preserving existing se
|
|
|
64
75
|
}
|
|
65
76
|
```
|
|
66
77
|
|
|
67
|
-
|
|
78
|
+
If an existing local bridge already imports `sortie-dogs/server`, `init` reuses it instead of adding
|
|
79
|
+
a duplicate package entry. A larger existing subagent depth is retained.
|
|
80
|
+
|
|
81
|
+
Completely restart OpenCode, then run:
|
|
68
82
|
|
|
69
83
|
```text
|
|
70
84
|
/sortie-v010 <task>
|
|
71
85
|
```
|
|
72
86
|
|
|
73
87
|
Selecting `dog-operator` directly starts the same workflow. `dog-operator` is the
|
|
74
|
-
|
|
88
|
+
user-facing entry point for the Mission profile. `dogs-coordinator` and every `*-v010` role are
|
|
75
89
|
internal children and must not be selected as task entry points.
|
|
76
90
|
|
|
77
|
-
`init` installs runtime assets and merges the required OpenCode settings; the
|
|
78
|
-
|
|
79
|
-
|
|
91
|
+
`init` installs runtime assets and merges the required OpenCode settings; the package entry or
|
|
92
|
+
existing local bridge loads enforcement and model routing. OpenCode can reload watched configuration,
|
|
93
|
+
but replacing an installed dependency may require a full restart. A new chat session alone does not
|
|
94
|
+
prove the newly installed plugin is loaded.
|
|
95
|
+
|
|
96
|
+
## v0.13.4 runtime updates
|
|
97
|
+
|
|
98
|
+
PR #152 restores same-session Coordinator/Operator implementation and formal validation through
|
|
99
|
+
`plan_units(executor="self")`, `start_direct_unit` and `finish_direct_unit`. Known single-unit work can
|
|
100
|
+
combine Mission start and planning; Luna Fast/max Worker routing remains the default. Full generated
|
|
101
|
+
contracts remain visible when an explicit Read line range covers the file. Native background
|
|
102
|
+
responsiveness remains: a launch acknowledgement or idle root is not Mission completion.
|
|
80
103
|
|
|
81
|
-
|
|
104
|
+
The initial independent Reviewer can investigate, correct, formally validate, deliver and self-recheck
|
|
105
|
+
continuously in its original native Task. Later Major/Medium findings accumulate in that same correction
|
|
106
|
+
context. Author self-recheck remains `self-rechecked`, `independent=false`, never independent `PASS`.
|
|
107
|
+
A different Reviewer is conditional on concrete residual Major risk; unresolved Major/Medium findings
|
|
108
|
+
still block acceptance. Operator owns final comparison and receipt.
|
|
82
109
|
|
|
83
|
-
|
|
110
|
+
Inherited compiler scratch no longer falsely invalidates broad-scope formal proof. Host-observed Git
|
|
111
|
+
delivery, caller-setting review and test-composition guidance reduce avoidable detours. Saved host
|
|
112
|
+
completion cards remain in tool history/UI, while outgoing V2 model/compaction context omits only their
|
|
113
|
+
presentation body, retaining receipt and evidence identities.
|
|
114
|
+
|
|
115
|
+
Two runs of the same fixed pre-release v8 package completed the original Anko task with official local
|
|
116
|
+
score 1 (F2P 9/9, P2P 94/94) and selected public probes 9/9. Times were 26m49s and 25m55s, costs
|
|
117
|
+
$1.47908048 and $1.54610816; each saved more than seven minutes versus the recorded v5 sample.
|
|
118
|
+
These limited same-task observations do not establish general speedup or a new SWE-bench score.
|
|
119
|
+
Release preflight, full tests, fixed-commit Windows CI and native Worker-start receipts are retained in
|
|
120
|
+
`_testenv/releases/0.13.4/`; startup/model identity is not task completion. See the
|
|
121
|
+
[release notes](docs/release-v0.13.4.md) and [quality-loop evidence](docs/nightly-quality-loop-20261002.md).
|
|
122
|
+
|
|
123
|
+
## Mission workflow
|
|
124
|
+
|
|
125
|
+
Use Operator → Worker when one useful unit and its meaningful formal check are known;
|
|
126
|
+
use Operator → Coordinator → Worker for actual discovery or decomposition:
|
|
84
127
|
|
|
85
128
|
- `dog-operator` states a few requirements/negative constraints and owns user decisions and final acceptance.
|
|
86
129
|
The host saves the original user message verbatim.
|
|
87
130
|
- Hidden `dogs-coordinator` owns investigation, unit declarations, Worker/Scout/Advisor/Reviewer dispatch,
|
|
88
131
|
in-request write-scope extensions, and corrections. It can read/search and run confirmation shell commands;
|
|
89
|
-
|
|
132
|
+
it can implement and formally validate directly in its own session, or delegate a unit to Worker.
|
|
90
133
|
- `dog-worker-v010` implements a host-generated unit within its file/directory write scopes.
|
|
91
134
|
Investigation commands need no pre-registration; formal checks retain real host-recorded results.
|
|
92
|
-
- High-risk changes require an independent Reviewer. Low-risk skips
|
|
93
|
-
|
|
135
|
+
- High-risk changes require an initial independent Reviewer with read/search access. Low-risk skips
|
|
136
|
+
are explicit and recorded. Reviewer-owned corrections follow the self-recheck policy above.
|
|
137
|
+
- Fast-lane can include high-risk single-unit work; it never implies a review skip.
|
|
138
|
+
- Investigation, edits, formal checks and requested Git delivery stay in the same implementing child.
|
|
139
|
+
Explicit user ordering is retained; no routine plan-approval or commit-only handoff is needed.
|
|
94
140
|
- Unit progress appears on the running Task without stopping Coordinator or prompting Operator.
|
|
95
141
|
|
|
96
|
-
The
|
|
97
|
-
parallel integration path are not exposed in this profile. More agents are not a
|
|
142
|
+
The `v010` Mission profile is serial by design; background responsiveness does not add parallel writers.
|
|
143
|
+
The stable profile's Luna fabric and parallel integration path are not exposed in this profile. More agents are not a
|
|
98
144
|
goal; preserving quality while reducing unnecessary expensive work is.
|
|
99
145
|
|
|
100
146
|
### SWE-bench evaluation
|
|
@@ -169,32 +215,51 @@ astroid **3/5**, pyvista **0/1**, sqlfluff **1/5**.
|
|
|
169
215
|
|
|
170
216
|
## Mission tools
|
|
171
217
|
|
|
172
|
-
1. `start_mission`: Operator supplies concise requirements; the host saves original messages
|
|
173
|
-
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
|
|
218
|
+
1. `start_mission`: Operator supplies concise requirements; the host saves original messages. A known single unit
|
|
219
|
+
can include `unit` to combine start/planning and return its configured Worker task.
|
|
220
|
+
2. `plan_units`: Operator or Coordinator supplies title, objective, file/directory scopes and formal checks. The host generates
|
|
221
|
+
IDs, handoff, manifest, proof mapping and the ready Worker task. `executor="self"` keeps execution in the
|
|
222
|
+
same controller session; `start_direct_unit` / `finish_direct_unit` retain observed formal-check freshness.
|
|
223
|
+
3. `operator_next`: advance serial units. `expand_unit` reconciles required in-request outputs while
|
|
224
|
+
preserving the same Task. A reasoned `plan_units` correction or `retry_mission_unit` handles ordinary
|
|
225
|
+
unit recovery under the original requirements and cumulative budget.
|
|
177
226
|
4. `review_mission`: generate the independent review packet from source, requirements and observed checks;
|
|
178
227
|
dispatch its Reviewer task for high-risk changes or record a low-risk skip.
|
|
179
|
-
5. `
|
|
180
|
-
|
|
228
|
+
5. `repair_review`: record findings and continue correction in the running original Reviewer Task, or resume
|
|
229
|
+
that same native session. Later findings accumulate; run inherited checks/requested Git delivery and
|
|
230
|
+
explicitly report `SELF_RECHECKED` in that same Task. Legacy
|
|
231
|
+
`CORRECTION_READY` alone requires a same-author read-only fallback through `review_mission`.
|
|
232
|
+
6. `submit_mission`: Coordinator returns a completion candidate, user-only decision, or proven external/scope/budget blocker.
|
|
233
|
+
7. `complete_mission`: Operator compares the original request, source and evidence, then explicitly accepts.
|
|
181
234
|
Only a succeeded receipt authorizes DONE and the measured 🐾 return report.
|
|
182
235
|
|
|
183
236
|
All tool names use the `sortie_v010_` prefix. Prior proposal/plan-repair tools remain in the compatibility
|
|
184
|
-
implementation but are hidden from the normal
|
|
185
|
-
|
|
237
|
+
implementation but are hidden from the normal Mission tool list. Operator/Coordinator handle in-request
|
|
238
|
+
path reconciliation without a new user approval. Changes beyond the original requirements or cumulative
|
|
239
|
+
budget return to Operator/user. `EVIDENCE_GAPS` is an advisory limitation, not Review `PASS` or an
|
|
240
|
+
automatic extra review; failed or missing required checks still prevent acceptance.
|
|
186
241
|
|
|
187
242
|
Durable profile state and hash-bound task references support restart and
|
|
188
243
|
compaction recovery without reconstructing criteria from summary prose. Stale,
|
|
189
|
-
foreign-root, or changed references are rejected.
|
|
190
|
-
|
|
191
|
-
|
|
244
|
+
foreign-root, or changed references are rejected. Requested `git add <paths>` and `git commit -m ...`
|
|
245
|
+
use the actual source write scope, not a fabricated `.git/**` scope. An optional host-managed Git
|
|
246
|
+
lifecycle also retains its branch, commit and post-commit boundaries. Neither mode grants arbitrary
|
|
247
|
+
Git, force push, release or publication authority.
|
|
248
|
+
|
|
249
|
+
### Progress and acceptance evidence
|
|
250
|
+
|
|
251
|
+
`sortie_v010_operator_status` keeps original requests, formal command/exit/timing observations,
|
|
252
|
+
review disposition and recorded delivery in a compact Mission view. `{ "view": "progress" }` exposes
|
|
253
|
+
the current unit, completed/total units, host budget and next action; `{ "view": "full" }` or
|
|
254
|
+
`details_ref` provides full snapshot diagnostics. Unknown clean state or a failed commit is not delivery
|
|
255
|
+
success. Reading progress does not dispatch, retry or accept work; use native completion notifications
|
|
256
|
+
instead of polling. Existing status reconciliation can recover a missed child terminal event.
|
|
192
257
|
|
|
193
258
|
## Configuration
|
|
194
259
|
|
|
195
260
|
### Profile files and precedence
|
|
196
261
|
|
|
197
|
-
The default package entry is the
|
|
262
|
+
The default package entry is the `v010` Mission profile:
|
|
198
263
|
|
|
199
264
|
- Command: `/sortie-v010`
|
|
200
265
|
- Primary agent: `dog-operator`
|
|
@@ -204,9 +269,11 @@ The default package entry is the v0.10 profile:
|
|
|
204
269
|
- Runtime state: `.sortie-dogs-v010/`
|
|
205
270
|
- Installed asset marker: `.opencode/sortie-dogs-v010.version`
|
|
206
271
|
|
|
272
|
+
For a global install, the marker is `<OpenCode config root>/sortie-dogs-v010.version`.
|
|
273
|
+
|
|
207
274
|
Precedence is built-in defaults, global file, project file, environment JSON,
|
|
208
275
|
then plugin factory options. Unknown properties or invalid types are rejected.
|
|
209
|
-
Use external
|
|
276
|
+
Use external `v010` role names such as `dog-operator`, `dogs-coordinator`, and
|
|
210
277
|
`dog-reviewer-v010` in `modelRouting`; do not also declare their stable aliases.
|
|
211
278
|
|
|
212
279
|
Example `.opencode/sortie-dogs-v010.json`:
|
|
@@ -249,8 +316,8 @@ not invent, probe, or translate variant names.
|
|
|
249
316
|
- `freeTierFallbackModels`: ordered global last-resort model IDs. Default:
|
|
250
317
|
`opencode/deepseek-v4-flash-free`; `[]` disables this fallback.
|
|
251
318
|
- `dedicatedWorkerModel`: canonical stable serial target, default
|
|
252
|
-
`openai/gpt-6.1-sol` / `medium`. The
|
|
253
|
-
role routes below; do not infer
|
|
319
|
+
`openai/gpt-6.1-sol` / `medium`. The Mission profile supplies its explicit
|
|
320
|
+
role routes below; do not infer its Worker route from this stable setting.
|
|
254
321
|
- `consultation.strategy`: fixed advisor identity, optional `required`, and
|
|
255
322
|
positive `maxCallsPerCandidate`; default one call and not required.
|
|
256
323
|
- `consultation.sourceReview`: risk-based review with `maxCallsPerCandidate`
|
|
@@ -264,27 +331,35 @@ not invent, probe, or translate variant names.
|
|
|
264
331
|
- `continuation.summarizeModel`: optional explicit compaction model; omission
|
|
265
332
|
reuses the latest observed root model.
|
|
266
333
|
- `validationProfile`: `fast`, `balanced`, or `assurance`; default `balanced`.
|
|
267
|
-
- `reflection`:
|
|
268
|
-
|
|
269
|
-
|
|
334
|
+
- `reflection`: enabled by default for `run`, `project` and `global` layers, with at most three
|
|
335
|
+
entries / 500 estimated tokens injected. Root Operator can use `sortie_v010_reflection` to retain
|
|
336
|
+
verified process causes/preventions for later turns and sessions; this is not model training.
|
|
337
|
+
Storage and managed blocks are separate from stable. Set `reflection.enabled` to `false` to disable.
|
|
270
338
|
|
|
271
|
-
The
|
|
339
|
+
The Mission host owns handoff and manifest controls under
|
|
272
340
|
`.sortie-dogs-v010/contracts/`. Do not create a legacy root
|
|
273
341
|
`operation-manifest.json` for this profile and do not edit generated controls.
|
|
274
342
|
Delete `.sortie-dogs-v010/` only when no Sortie run is active.
|
|
275
343
|
|
|
276
344
|
### Validation policy
|
|
277
345
|
|
|
278
|
-
`validationProfile` chooses non-canonical depth:
|
|
346
|
+
`validationProfile` chooses supplementary non-canonical depth:
|
|
279
347
|
|
|
280
348
|
- `fast`: static checks
|
|
281
349
|
- `balanced`: targeted checks
|
|
282
350
|
- `assurance`: related checks
|
|
283
351
|
|
|
284
|
-
|
|
285
|
-
|
|
286
|
-
|
|
287
|
-
|
|
352
|
+
It does not replace meaningful declared formal checks or user/project-required broad validation.
|
|
353
|
+
Batch related edits, run focused checks, then execute every declared formal check in order on the
|
|
354
|
+
stable candidate. The implementing Worker or admitted correcting Reviewer runs those commands;
|
|
355
|
+
canonical/full-suite `owner=coordinator` is evidence accounting, not a requirement for root execution.
|
|
356
|
+
|
|
357
|
+
Keep required broad checks for the final integrated candidate. Reuse valid unchanged evidence only
|
|
358
|
+
when the contract permits; identity includes candidate, command, environment, scope and owner.
|
|
359
|
+
Required repeated occurrences retain their own execution identities and cannot be skipped as duplicates.
|
|
360
|
+
Native commands, actual working directory, exit, duration and saved source bindings establish freshness.
|
|
361
|
+
Diagnostics do not substitute for formal proof. Repeat checks when changes, failures or freshness require
|
|
362
|
+
it, rather than solely because a Worker changed or documentation was edited.
|
|
288
363
|
|
|
289
364
|
### Default routes
|
|
290
365
|
|
|
@@ -315,25 +390,40 @@ export { SortieDogsPlugin } from "sortie-dogs/plugin/stable";
|
|
|
315
390
|
|
|
316
391
|
The stable profile uses `/sortie`, `dog-coordinator`,
|
|
317
392
|
`.opencode/sortie-dogs.json`, `SORTIE_DOGS_CONFIG`, and `.sortie-dogs/`. Do not
|
|
318
|
-
register stable and
|
|
393
|
+
register stable and `v010` from the same package installation path in one host.
|
|
319
394
|
|
|
320
395
|
## Global availability
|
|
321
396
|
|
|
322
|
-
Project-local installation is recommended. To expose
|
|
397
|
+
Project-local installation is recommended. To expose the current Mission assets globally:
|
|
323
398
|
|
|
324
399
|
```sh
|
|
325
|
-
npm install --global sortie-dogs
|
|
400
|
+
npm install --global sortie-dogs@0.13.4
|
|
326
401
|
sortie-dogs init --global --profile v010
|
|
327
402
|
```
|
|
328
403
|
|
|
329
|
-
Global initialization
|
|
330
|
-
|
|
404
|
+
Global initialization registers the package or reuses an existing local V2 bridge, and sets subagent
|
|
405
|
+
depth to at least two, preserving unrelated settings and the default agent.
|
|
406
|
+
|
|
407
|
+
An existing `<OpenCode config root>/plugins/sortie-dogs/index.js` bridge importing `sortie-dogs/server`
|
|
408
|
+
can resolve a **separate dependency** under that config root. Updating npm-global alone does not update
|
|
409
|
+
it. For that layout, also install the same release at the actual config root, then rerun global init:
|
|
410
|
+
|
|
411
|
+
```sh
|
|
412
|
+
npm install --prefix "$HOME/.config/opencode" sortie-dogs@0.13.4
|
|
413
|
+
sortie-dogs init --global --profile v010
|
|
414
|
+
```
|
|
415
|
+
|
|
416
|
+
The command shows the default config root; use your actual root if overridden. For a configured npm
|
|
417
|
+
package entry, OpenCode V2 also provides `opencode plugin list` / `opencode plugin update`; exact
|
|
418
|
+
version pins require an explicit version change. Completely restart OpenCode after updating, then
|
|
419
|
+
check the installed package, asset marker and loaded plugin version.
|
|
331
420
|
|
|
332
421
|
## Updates and removal
|
|
333
422
|
|
|
334
|
-
|
|
423
|
+
For project-local updates, replace the dependency, rerun initialization and completely restart OpenCode:
|
|
335
424
|
|
|
336
425
|
```sh
|
|
426
|
+
npm install --save-dev sortie-dogs@latest
|
|
337
427
|
npx sortie-dogs init .
|
|
338
428
|
```
|
|
339
429
|
|
|
@@ -341,6 +431,10 @@ npx sortie-dogs init .
|
|
|
341
431
|
version, preserves user configuration, and stops safely on unknown ownership or
|
|
342
432
|
conflicting files.
|
|
343
433
|
|
|
434
|
+
Align any exact version pin or separate bridge dependency with the intended release too. An installed
|
|
435
|
+
marker of `0.13.4-reviewer-continuous-v1` identifies the assets; it does not prove an already-running
|
|
436
|
+
OpenCode process has reloaded the plugin.
|
|
437
|
+
|
|
344
438
|
There is no supported uninstall command. Remove the npm dependency separately,
|
|
345
439
|
then follow the [safe manual removal guide](docs/uninstall.md). Delete only known
|
|
346
440
|
Sortie-owned paths; never remove the whole `.opencode` directory or use broad
|
package/dist/asset-version.d.ts
CHANGED
|
@@ -3,5 +3,5 @@
|
|
|
3
3
|
* installed project marker without importing every asset body.
|
|
4
4
|
*/
|
|
5
5
|
export declare const RUNTIME_ASSET_VERSION = "0.3.89-completion-proof-v1";
|
|
6
|
-
export declare const V010_RUNTIME_ASSET_VERSION = "0.13.
|
|
6
|
+
export declare const V010_RUNTIME_ASSET_VERSION = "0.13.4-reviewer-continuous-v1";
|
|
7
7
|
export type RuntimeAssetVersion = typeof RUNTIME_ASSET_VERSION | typeof V010_RUNTIME_ASSET_VERSION;
|
package/dist/asset-version.js
CHANGED
|
@@ -3,4 +3,4 @@
|
|
|
3
3
|
* installed project marker without importing every asset body.
|
|
4
4
|
*/
|
|
5
5
|
export const RUNTIME_ASSET_VERSION = "0.3.89-completion-proof-v1";
|
|
6
|
-
export const V010_RUNTIME_ASSET_VERSION = "0.13.
|
|
6
|
+
export const V010_RUNTIME_ASSET_VERSION = "0.13.4-reviewer-continuous-v1";
|
|
@@ -6,3 +6,9 @@ export declare const CONTRACT_TEXT_LIMITS: Readonly<{
|
|
|
6
6
|
command: 8192;
|
|
7
7
|
path: 512;
|
|
8
8
|
}>;
|
|
9
|
+
/** New Mission task generation only; persisted/legacy contracts retain their original bounds.
|
|
10
|
+
* Only target is authoring guidance. maximum is internal headroom, never a new target. */
|
|
11
|
+
export declare const MISSION_OBJECTIVE_LIMITS: Readonly<{
|
|
12
|
+
target: 2000;
|
|
13
|
+
maximum: 3000;
|
|
14
|
+
}>;
|
|
@@ -1,2 +1,5 @@
|
|
|
1
1
|
/** Common text bounds shared by handoff, manifest and goal evidence validation. */
|
|
2
2
|
export const CONTRACT_TEXT_LIMITS = Object.freeze({ title: 160, objective: 32768, statement: 1000, command: 8192, path: 512 });
|
|
3
|
+
/** New Mission task generation only; persisted/legacy contracts retain their original bounds.
|
|
4
|
+
* Only target is authoring guidance. maximum is internal headroom, never a new target. */
|
|
5
|
+
export const MISSION_OBJECTIVE_LIMITS = Object.freeze({ target: 2000, maximum: 3000 });
|
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
import { type OperatorPlan, type OperatorState, type OperatorTask } from "./operator-runtime.js";
|
|
1
|
+
import { type OperatorPlan, type OperatorState, type OperatorTask, type OperatorRuntime } from "./operator-runtime.js";
|
|
2
2
|
import { type RuntimeProfile } from "./runtime-profile.js";
|
|
3
3
|
export declare const MISSION_REFERENCE = "SORTIE_MISSION_REF ";
|
|
4
4
|
export declare const MISSION_REVIEW_REFERENCE = "SORTIE_MISSION_REVIEW_REF ";
|
|
@@ -44,7 +44,7 @@ export interface MissionAttempt {
|
|
|
44
44
|
predecessorAttemptID?: string | null;
|
|
45
45
|
/** Fingerprint of the settled scoped candidate against which a later Rescue is proposed. */
|
|
46
46
|
candidateID?: string;
|
|
47
|
-
kind: "implementation" | "normal_remediation" | "astra_rescue";
|
|
47
|
+
kind: "implementation" | "normal_remediation" | "astra_rescue" | "reviewer_correction" | "direct_execution";
|
|
48
48
|
status: "pending" | "dispatched" | "succeeded" | "failed" | "cancelled" | "unconfirmed";
|
|
49
49
|
callID?: string;
|
|
50
50
|
childSessionID?: string;
|
|
@@ -100,6 +100,24 @@ export interface MissionLaunchConditions {
|
|
|
100
100
|
applies_to: string;
|
|
101
101
|
recordedAt?: string;
|
|
102
102
|
}
|
|
103
|
+
/** Native terminal self-recheck by the correction author, explicitly not independent approval. */
|
|
104
|
+
export interface MissionSelfRecheck {
|
|
105
|
+
runID: string;
|
|
106
|
+
source: string;
|
|
107
|
+
/** Candidate source/check identity, excluding optional excerpt presentation. */
|
|
108
|
+
candidateSource?: string;
|
|
109
|
+
author: string;
|
|
110
|
+
callID: string;
|
|
111
|
+
promptID: string;
|
|
112
|
+
messageID: string;
|
|
113
|
+
nativeOutcome: "completed";
|
|
114
|
+
result: string;
|
|
115
|
+
unresolvedFindings: string[];
|
|
116
|
+
residualMajor?: {
|
|
117
|
+
reachable_path: string;
|
|
118
|
+
consequence: string;
|
|
119
|
+
};
|
|
120
|
+
}
|
|
103
121
|
export interface OperatorMission {
|
|
104
122
|
version: "0.12";
|
|
105
123
|
id: string;
|
|
@@ -107,6 +125,12 @@ export interface OperatorMission {
|
|
|
107
125
|
requests: MissionRequest[];
|
|
108
126
|
/** Prior public conversation context, not additional immutable requirements. */
|
|
109
127
|
context?: MissionContext[];
|
|
128
|
+
/** Explicit continue deliveries, keyed by the original real turn. */
|
|
129
|
+
steering?: {
|
|
130
|
+
requestID: string;
|
|
131
|
+
child: string;
|
|
132
|
+
status: "pending" | "queued";
|
|
133
|
+
}[];
|
|
110
134
|
kind?: "implementation" | "operation";
|
|
111
135
|
/** Native shell observations of the requested operation, separate from auxiliary checks. */
|
|
112
136
|
execution?: MissionExecution;
|
|
@@ -127,6 +151,15 @@ export interface OperatorMission {
|
|
|
127
151
|
runID: string | null;
|
|
128
152
|
/** Git HEAD before this mission's first implementation unit, retained across replans and commits. */
|
|
129
153
|
reviewBaseline?: string;
|
|
154
|
+
/** A dated host Git observation, never an inference from a Worker report or lifecycle plan. */
|
|
155
|
+
deliveryObservation?: {
|
|
156
|
+
run_id: string;
|
|
157
|
+
observed_at: string;
|
|
158
|
+
head: string | null;
|
|
159
|
+
branch: string | null;
|
|
160
|
+
clean: boolean;
|
|
161
|
+
source: string;
|
|
162
|
+
};
|
|
130
163
|
/** Cumulative declared review inputs/outputs, including units that failed after writing source. */
|
|
131
164
|
reviewScope?: MissionReviewScope;
|
|
132
165
|
/** Optional Advisor/Scout decisions and native consultation outcomes; never an admission gate. */
|
|
@@ -152,17 +185,43 @@ export interface OperatorMission {
|
|
|
152
185
|
runID: string;
|
|
153
186
|
risk: string[];
|
|
154
187
|
source: string;
|
|
188
|
+
candidateSource?: string;
|
|
155
189
|
task: OperatorTask | null;
|
|
190
|
+
callID?: string;
|
|
156
191
|
evidence?: MissionEvidenceExcerpt[];
|
|
157
192
|
requestFingerprint?: string;
|
|
158
|
-
verdict: "pending" | "PASS" | "findings" | "evidence-gaps" | "skipped-low-risk";
|
|
193
|
+
verdict: "pending" | "PASS" | "findings" | "evidence-gaps" | "skipped-low-risk" | "self-rechecked";
|
|
159
194
|
result?: string;
|
|
160
195
|
child?: string;
|
|
196
|
+
mode?: "independent" | "self-recheck";
|
|
197
|
+
admittedAt?: number;
|
|
198
|
+
promptID?: string;
|
|
199
|
+
selfRecheck?: MissionSelfRecheck;
|
|
161
200
|
/** Completed, independent initial review for this mission, not merely an inherited child ID. */
|
|
162
201
|
initialPrompt?: string;
|
|
163
202
|
/** Observed evidence-only reviews; reporting only, never an acceptance threshold. */
|
|
164
203
|
evidenceGapReviews?: number;
|
|
165
204
|
};
|
|
205
|
+
/** Retained independently of later review generations; authors never become independent reviewers. */
|
|
206
|
+
corrections?: {
|
|
207
|
+
author: string;
|
|
208
|
+
reviewIdentity: string;
|
|
209
|
+
priorRunID: string;
|
|
210
|
+
runID: string;
|
|
211
|
+
priorSource: string;
|
|
212
|
+
findings: string;
|
|
213
|
+
initialPrompt: string;
|
|
214
|
+
baseline?: string;
|
|
215
|
+
/** Actual still-running initial Review; direct correction is not another child terminal. */
|
|
216
|
+
inlineReview?: {
|
|
217
|
+
callID: string;
|
|
218
|
+
promptID: string;
|
|
219
|
+
admittedAt: number;
|
|
220
|
+
reviewIdentity: string;
|
|
221
|
+
};
|
|
222
|
+
status: "prepared" | "running" | "ready" | "failed" | "cancelled";
|
|
223
|
+
selfRecheck?: MissionSelfRecheck;
|
|
224
|
+
}[];
|
|
166
225
|
}
|
|
167
226
|
export declare const MISSION_CONSULTATION_LIMIT = 32;
|
|
168
227
|
/** Review coverage survives a narrower replan; it is not a Worker write grant. */
|
|
@@ -171,6 +230,8 @@ export declare function missionReviewScope(previous: MissionReviewScope | undefi
|
|
|
171
230
|
export declare function missionReviewVerdict(text: string): "PASS" | "evidence-gaps" | "findings";
|
|
172
231
|
/** Whether the recorded review permits submission and acceptance of the current candidate. */
|
|
173
232
|
export declare function missionReviewAccepted(review: NonNullable<OperatorMission["review"]>): boolean;
|
|
233
|
+
/** A short native report, not a tag/hash-based second-review policy or an approval checklist. */
|
|
234
|
+
export declare function missionSelfRecheckReport(text: string, source: string, hostBound?: boolean): Pick<MissionSelfRecheck, "unresolvedFindings" | "residualMajor"> | undefined;
|
|
174
235
|
export declare function missionExecutionStatus(mission: OperatorMission): "not-required" | "not-started" | "running" | "execution-failed" | "executed";
|
|
175
236
|
/** Observe the command's own terminal summary, never Worker/Reviewer prose. Scores are result data,
|
|
176
237
|
* not process success. Unstructured commands retain their native exit semantics. */
|
|
@@ -187,6 +248,7 @@ export declare class OperatorMissionRuntime {
|
|
|
187
248
|
private readonly writes;
|
|
188
249
|
constructor(projectRoot: string, profile: RuntimeProfile);
|
|
189
250
|
private file;
|
|
251
|
+
correctionReference(root: string, reviewIdentity: string): string;
|
|
190
252
|
private load;
|
|
191
253
|
/** Recover non-replacement continuations only; an explicit replacement must link the current cancelled run. */
|
|
192
254
|
private loadMission;
|
|
@@ -228,7 +290,10 @@ export declare class OperatorMissionRuntime {
|
|
|
228
290
|
brief(state: OperatorMission): string;
|
|
229
291
|
}
|
|
230
292
|
/** The model supplies only useful unit facts; IDs, proof projection and control documents are generated here. */
|
|
231
|
-
export declare function missionPlan(mission: OperatorMission, raw: unknown): OperatorPlan;
|
|
293
|
+
export declare function missionPlan(mission: OperatorMission, raw: unknown, projectRoot?: string): OperatorPlan;
|
|
294
|
+
/** Compact provenance, not a new acceptance verdict or a substitute for original-request comparison. */
|
|
295
|
+
export declare function missionAcceptanceSummary(mission: OperatorMission, run: OperatorState | undefined, operators: OperatorRuntime, observe?: (validation: readonly string[], child: string | null, notBefore?: number) => Promise<unknown>): Promise<Record<string, unknown>>;
|
|
296
|
+
export declare function missionReviewIndependent(mission: OperatorMission, child: string | undefined): boolean;
|
|
232
297
|
export declare function missionPacket(mission: OperatorMission, run?: OperatorState): Record<string, unknown>;
|
|
233
298
|
/** Models forward a short capability, never recopy the host's evidence hashes and source packet. */
|
|
234
299
|
export declare function missionReviewTask(mission: OperatorMission): OperatorTask;
|