sortie-dogs 0.13.2 → 0.13.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (39) hide show
  1. package/README.md +142 -48
  2. package/dist/asset-version.d.ts +1 -1
  3. package/dist/asset-version.js +1 -1
  4. package/dist/core/contract-limits.d.ts +6 -0
  5. package/dist/core/contract-limits.js +3 -0
  6. package/dist/core/operator-mission.d.ts +69 -4
  7. package/dist/core/operator-mission.js +148 -31
  8. package/dist/core/operator-runtime.d.ts +36 -0
  9. package/dist/core/operator-runtime.js +189 -18
  10. package/dist/core/runtime-profile.js +3 -0
  11. package/dist/core/validate-schema.js +0 -1
  12. package/dist/core/validation-budget.d.ts +5 -0
  13. package/dist/core/validation-budget.js +5 -2
  14. package/dist/plugin/gate.d.ts +7 -1
  15. package/dist/plugin/gate.js +131 -24
  16. package/dist/plugin/index.d.ts +17 -0
  17. package/dist/plugin/index.js +354 -59
  18. package/dist/plugin/mission-review.d.ts +112 -2
  19. package/dist/plugin/mission-review.js +225 -36
  20. package/dist/plugin/native-background.d.ts +40 -0
  21. package/dist/plugin/native-background.js +211 -0
  22. package/dist/plugin/native-contract-read.d.ts +9 -0
  23. package/dist/plugin/native-contract-read.js +98 -0
  24. package/dist/plugin/profiled.js +936 -182
  25. package/dist/plugin/protected-snapshot.d.ts +4 -0
  26. package/dist/plugin/protected-snapshot.js +51 -13
  27. package/dist/plugin/receipt-presentation.d.ts +2 -0
  28. package/dist/plugin/receipt-presentation.js +48 -0
  29. package/dist/plugin/run-metrics.d.ts +1 -1
  30. package/dist/plugin/run-metrics.js +3 -2
  31. package/dist/plugin/runtime-bridge.d.ts +67 -2
  32. package/dist/plugin/v2.d.ts +40 -20
  33. package/dist/plugin/v2.js +347 -55
  34. package/dist/plugin/validation-scratch.d.ts +2 -2
  35. package/dist/plugin/validation-scratch.js +29 -17
  36. package/dist/runtime-assets-v010.js +14 -6
  37. package/dist/runtime-mission-assets.d.ts +5 -3
  38. package/dist/runtime-mission-assets.js +188 -133
  39. package/package.json +1 -1
package/README.md CHANGED
@@ -14,6 +14,9 @@ implementation, validation, review, and model routing.
14
14
  continuation, remediation, and restart.
15
15
  - **Adaptive execution**: small work stays small; additional agents and stronger
16
16
  models are used only when task shape or risk justifies them.
17
+ - **Clear Worker instructions**: give Luna a concise goal, explicit constraints
18
+ and completion criteria; keep orchestration bookkeeping in the harness.
19
+ See the [instruction design principles](docs/worker-instruction-design.md).
17
20
  - **Coexistence**: Sortie activates only when selected and preserves normal
18
21
  OpenCode agents, settings, and user-owned files.
19
22
  - **Cost, time, and proof**: the objective is a verified result at the lowest
@@ -32,6 +35,11 @@ implementation, validation, review, and model routing.
32
35
  Guides: [日本語](docs/guide-ja.md) · [简体中文](docs/guide-zh-CN.md) ·
33
36
  [Testing](docs/testing.md) · [CLI testing](docs/cli-testing.md)
34
37
 
38
+ **Current release: [v0.13.4](https://github.com/zufall-upon/Sortie-dogs/releases/tag/v0.13.4)**
39
+ ([release notes](docs/release-v0.13.4.md)). The default Mission runtime retains the `v010`
40
+ profile, command and configuration names for compatibility; these names do not mean v0.10 is installed.
41
+ The current asset marker is `0.13.4-reviewer-continuous-v1`.
42
+
35
43
  ## SWE-bench Lite: 170/300 (56.67%)
36
44
 
37
45
  The fixed **Sortie-dogs v0.12.24** harness resolved **170 of 300 SWE-bench Lite test issues** in one pass@1 campaign, with 9 empty patches and no official evaluation errors. Every instance has a frozen prediction and an inference-time trajectory. The task Workers ran `openai/gpt-6-luna-fast#max`; operator, coordinator and review roles ran `openai/gpt-6-sol#xhigh`. This is a system result, **not** a Luna-only model comparison or a Verified/full SWE-bench score.
@@ -40,22 +48,25 @@ The fixed **Sortie-dogs v0.12.24** harness resolved **170 of 300 SWE-bench Lite
40
48
 
41
49
  The single official 300-instance report and frozen predictions are hash-bound in the report. Confirmed inference expense was **$162.99**; a separate **$34.60** of usage has unknown pricing and is held against the campaign cap, **not** counted as known expense. Leaderboard registration and maintainer acceptance are separate from this official local evaluation.
42
50
 
43
- > **Beta:** v0.12.2 builds on the v0.10.23 execution engine. Runtime behavior,
51
+ Historical scores below belong to their fixed candidates, not v0.13.4. SWE-bench is a separate,
52
+ optional measurement rather than a mandatory release gate.
53
+
54
+ > **Beta:** v0.13.x is still stabilizing. Runtime behavior,
44
55
  > configuration, and generated assets may still change before 1.0.
45
56
 
46
57
  ## Quick start
47
58
 
48
- Requirements: Node.js 22.6 or newer, npm, and OpenCode.
59
+ Requirements: Node.js 22.6 or newer, npm, and OpenCode V2.
49
60
 
50
61
  Run these commands in the target project:
51
62
 
52
63
  ```sh
53
- npm install --save-dev sortie-dogs
64
+ npm install --save-dev sortie-dogs@latest
54
65
  npx sortie-dogs init .
55
66
  ```
56
67
 
57
- The package retains the `v010` profile/namespace for compatibility. `init` adds the OpenCode V2 plugin and the required
58
- two-level subagent depth to `.opencode/opencode.json(c)`, preserving existing settings:
68
+ `init` defaults to the `v010` Mission profile and registers the OpenCode V2 plugin in
69
+ `.opencode/opencode.json(c)`, preserving unrelated settings. It sets subagent depth to at least two:
59
70
 
60
71
  ```json
61
72
  {
@@ -64,37 +75,72 @@ two-level subagent depth to `.opencode/opencode.json(c)`, preserving existing se
64
75
  }
65
76
  ```
66
77
 
67
- Restart OpenCode, then run:
78
+ If an existing local bridge already imports `sortie-dogs/server`, `init` reuses it instead of adding
79
+ a duplicate package entry. A larger existing subagent depth is retained.
80
+
81
+ Completely restart OpenCode, then run:
68
82
 
69
83
  ```text
70
84
  /sortie-v010 <task>
71
85
  ```
72
86
 
73
87
  Selecting `dog-operator` directly starts the same workflow. `dog-operator` is the
74
- only user-facing v0.10 authority. `dogs-coordinator` and every `*-v010` role are
88
+ user-facing entry point for the Mission profile. `dogs-coordinator` and every `*-v010` role are
75
89
  internal children and must not be selected as task entry points.
76
90
 
77
- `init` installs runtime assets and merges the required OpenCode settings; the `plugins` entry loads runtime enforcement and
78
- model routing. A new session alone does not reload an updated
79
- plugin process, so restart OpenCode after installation or upgrade.
91
+ `init` installs runtime assets and merges the required OpenCode settings; the package entry or
92
+ existing local bridge loads enforcement and model routing. OpenCode can reload watched configuration,
93
+ but replacing an installed dependency may require a full restart. A new chat session alone does not
94
+ prove the newly installed plugin is loaded.
95
+
96
+ ## v0.13.4 runtime updates
97
+
98
+ PR #152 restores same-session Coordinator/Operator implementation and formal validation through
99
+ `plan_units(executor="self")`, `start_direct_unit` and `finish_direct_unit`. Known single-unit work can
100
+ combine Mission start and planning; Luna Fast/max Worker routing remains the default. Full generated
101
+ contracts remain visible when an explicit Read line range covers the file. Native background
102
+ responsiveness remains: a launch acknowledgement or idle root is not Mission completion.
80
103
 
81
- ## v0.12.0 workflow
104
+ The initial independent Reviewer can investigate, correct, formally validate, deliver and self-recheck
105
+ continuously in its original native Task. Later Major/Medium findings accumulate in that same correction
106
+ context. Author self-recheck remains `self-rechecked`, `independent=false`, never independent `PASS`.
107
+ A different Reviewer is conditional on concrete residual Major risk; unresolved Major/Medium findings
108
+ still block acceptance. Operator owns final comparison and receipt.
82
109
 
83
- v0.12.0 keeps Operator → Coordinator → Worker, with independent review:
110
+ Inherited compiler scratch no longer falsely invalidates broad-scope formal proof. Host-observed Git
111
+ delivery, caller-setting review and test-composition guidance reduce avoidable detours. Saved host
112
+ completion cards remain in tool history/UI, while outgoing V2 model/compaction context omits only their
113
+ presentation body, retaining receipt and evidence identities.
114
+
115
+ Two runs of the same fixed pre-release v8 package completed the original Anko task with official local
116
+ score 1 (F2P 9/9, P2P 94/94) and selected public probes 9/9. Times were 26m49s and 25m55s, costs
117
+ $1.47908048 and $1.54610816; each saved more than seven minutes versus the recorded v5 sample.
118
+ These limited same-task observations do not establish general speedup or a new SWE-bench score.
119
+ Release preflight, full tests, fixed-commit Windows CI and native Worker-start receipts are retained in
120
+ `_testenv/releases/0.13.4/`; startup/model identity is not task completion. See the
121
+ [release notes](docs/release-v0.13.4.md) and [quality-loop evidence](docs/nightly-quality-loop-20261002.md).
122
+
123
+ ## Mission workflow
124
+
125
+ Use Operator → Worker when one useful unit and its meaningful formal check are known;
126
+ use Operator → Coordinator → Worker for actual discovery or decomposition:
84
127
 
85
128
  - `dog-operator` states a few requirements/negative constraints and owns user decisions and final acceptance.
86
129
  The host saves the original user message verbatim.
87
130
  - Hidden `dogs-coordinator` owns investigation, unit declarations, Worker/Scout/Advisor/Reviewer dispatch,
88
131
  in-request write-scope extensions, and corrections. It can read/search and run confirmation shell commands;
89
- source editing tools belong to Worker.
132
+ it can implement and formally validate directly in its own session, or delegate a unit to Worker.
90
133
  - `dog-worker-v010` implements a host-generated unit within its file/directory write scopes.
91
134
  Investigation commands need no pre-registration; formal checks retain real host-recorded results.
92
- - High-risk changes require an independent Reviewer. Low-risk skips are explicit and recorded.
93
- - Simple low-risk single-unit work retains Operator → Worker Fast-lane dispatch.
135
+ - High-risk changes require an initial independent Reviewer with read/search access. Low-risk skips
136
+ are explicit and recorded. Reviewer-owned corrections follow the self-recheck policy above.
137
+ - Fast-lane can include high-risk single-unit work; it never implies a review skip.
138
+ - Investigation, edits, formal checks and requested Git delivery stay in the same implementing child.
139
+ Explicit user ordering is retained; no routine plan-approval or commit-only handoff is needed.
94
140
  - Unit progress appears on the running Task without stopping Coordinator or prompting Operator.
95
141
 
96
- The v0.10 profile is serial by design. The stable profile's Luna fabric and
97
- parallel integration path are not exposed in this profile. More agents are not a
142
+ The `v010` Mission profile is serial by design; background responsiveness does not add parallel writers.
143
+ The stable profile's Luna fabric and parallel integration path are not exposed in this profile. More agents are not a
98
144
  goal; preserving quality while reducing unnecessary expensive work is.
99
145
 
100
146
  ### SWE-bench evaluation
@@ -169,32 +215,51 @@ astroid **3/5**, pyvista **0/1**, sqlfluff **1/5**.
169
215
 
170
216
  ## Mission tools
171
217
 
172
- 1. `start_mission`: Operator supplies concise requirements; the host saves original messages and returns a Coordinator task.
173
- 2. `plan_units`: Coordinator supplies title, objective, file/directory scopes and formal checks. The host generates
174
- IDs, handoff, manifest, proof mapping and the ready Worker task. No proposal approval round trip.
175
- 3. `operator_next`: advance serial units. `expand_unit` or a reasoned `plan_units` correction extends/replaces
176
- settled execution within the original requirements and retained cumulative budget.
218
+ 1. `start_mission`: Operator supplies concise requirements; the host saves original messages. A known single unit
219
+ can include `unit` to combine start/planning and return its configured Worker task.
220
+ 2. `plan_units`: Operator or Coordinator supplies title, objective, file/directory scopes and formal checks. The host generates
221
+ IDs, handoff, manifest, proof mapping and the ready Worker task. `executor="self"` keeps execution in the
222
+ same controller session; `start_direct_unit` / `finish_direct_unit` retain observed formal-check freshness.
223
+ 3. `operator_next`: advance serial units. `expand_unit` reconciles required in-request outputs while
224
+ preserving the same Task. A reasoned `plan_units` correction or `retry_mission_unit` handles ordinary
225
+ unit recovery under the original requirements and cumulative budget.
177
226
  4. `review_mission`: generate the independent review packet from source, requirements and observed checks;
178
227
  dispatch its Reviewer task for high-risk changes or record a low-risk skip.
179
- 5. `submit_mission`: Coordinator returns a completion candidate, user-only decision, or proven external/scope/budget blocker.
180
- 6. `complete_mission`: Operator compares the original request, source and evidence, then explicitly accepts.
228
+ 5. `repair_review`: record findings and continue correction in the running original Reviewer Task, or resume
229
+ that same native session. Later findings accumulate; run inherited checks/requested Git delivery and
230
+ explicitly report `SELF_RECHECKED` in that same Task. Legacy
231
+ `CORRECTION_READY` alone requires a same-author read-only fallback through `review_mission`.
232
+ 6. `submit_mission`: Coordinator returns a completion candidate, user-only decision, or proven external/scope/budget blocker.
233
+ 7. `complete_mission`: Operator compares the original request, source and evidence, then explicitly accepts.
181
234
  Only a succeeded receipt authorizes DONE and the measured 🐾 return report.
182
235
 
183
236
  All tool names use the `sortie_v010_` prefix. Prior proposal/plan-repair tools remain in the compatibility
184
- implementation but are hidden from the normal v0.12 model tool list. An in-request path extension is a
185
- Coordinator decision; changing the original requirements or increasing budget returns to Operator/user.
237
+ implementation but are hidden from the normal Mission tool list. Operator/Coordinator handle in-request
238
+ path reconciliation without a new user approval. Changes beyond the original requirements or cumulative
239
+ budget return to Operator/user. `EVIDENCE_GAPS` is an advisory limitation, not Review `PASS` or an
240
+ automatic extra review; failed or missing required checks still prevent acceptance.
186
241
 
187
242
  Durable profile state and hash-bound task references support restart and
188
243
  compaction recovery without reconstructing criteria from summary prose. Stale,
189
- foreign-root, or changed references are rejected. An optional Git lifecycle can
190
- create one non-overwriting branch and one explicit-path commit; it never grants
191
- arbitrary Git, force push, release, or publication authority.
244
+ foreign-root, or changed references are rejected. Requested `git add <paths>` and `git commit -m ...`
245
+ use the actual source write scope, not a fabricated `.git/**` scope. An optional host-managed Git
246
+ lifecycle also retains its branch, commit and post-commit boundaries. Neither mode grants arbitrary
247
+ Git, force push, release or publication authority.
248
+
249
+ ### Progress and acceptance evidence
250
+
251
+ `sortie_v010_operator_status` keeps original requests, formal command/exit/timing observations,
252
+ review disposition and recorded delivery in a compact Mission view. `{ "view": "progress" }` exposes
253
+ the current unit, completed/total units, host budget and next action; `{ "view": "full" }` or
254
+ `details_ref` provides full snapshot diagnostics. Unknown clean state or a failed commit is not delivery
255
+ success. Reading progress does not dispatch, retry or accept work; use native completion notifications
256
+ instead of polling. Existing status reconciliation can recover a missed child terminal event.
192
257
 
193
258
  ## Configuration
194
259
 
195
260
  ### Profile files and precedence
196
261
 
197
- The default package entry is the v0.10 profile:
262
+ The default package entry is the `v010` Mission profile:
198
263
 
199
264
  - Command: `/sortie-v010`
200
265
  - Primary agent: `dog-operator`
@@ -204,9 +269,11 @@ The default package entry is the v0.10 profile:
204
269
  - Runtime state: `.sortie-dogs-v010/`
205
270
  - Installed asset marker: `.opencode/sortie-dogs-v010.version`
206
271
 
272
+ For a global install, the marker is `<OpenCode config root>/sortie-dogs-v010.version`.
273
+
207
274
  Precedence is built-in defaults, global file, project file, environment JSON,
208
275
  then plugin factory options. Unknown properties or invalid types are rejected.
209
- Use external v0.10 role names such as `dog-operator`, `dogs-coordinator`, and
276
+ Use external `v010` role names such as `dog-operator`, `dogs-coordinator`, and
210
277
  `dog-reviewer-v010` in `modelRouting`; do not also declare their stable aliases.
211
278
 
212
279
  Example `.opencode/sortie-dogs-v010.json`:
@@ -249,8 +316,8 @@ not invent, probe, or translate variant names.
249
316
  - `freeTierFallbackModels`: ordered global last-resort model IDs. Default:
250
317
  `opencode/deepseek-v4-flash-free`; `[]` disables this fallback.
251
318
  - `dedicatedWorkerModel`: canonical stable serial target, default
252
- `openai/gpt-6.1-sol` / `medium`. The v0.10 profile also supplies its explicit
253
- role routes below; do not infer the v0.10 worker route from this stable setting.
319
+ `openai/gpt-6.1-sol` / `medium`. The Mission profile supplies its explicit
320
+ role routes below; do not infer its Worker route from this stable setting.
254
321
  - `consultation.strategy`: fixed advisor identity, optional `required`, and
255
322
  positive `maxCallsPerCandidate`; default one call and not required.
256
323
  - `consultation.sourceReview`: risk-based review with `maxCallsPerCandidate`
@@ -264,27 +331,35 @@ not invent, probe, or translate variant names.
264
331
  - `continuation.summarizeModel`: optional explicit compaction model; omission
265
332
  reuses the latest observed root model.
266
333
  - `validationProfile`: `fast`, `balanced`, or `assurance`; default `balanced`.
267
- - `reflection`: accepted by the shared schema, but reflection writes are not
268
- exposed by the serial v0.10 profile. Stable reflection remains opt-in and off by
269
- default.
334
+ - `reflection`: enabled by default for `run`, `project` and `global` layers, with at most three
335
+ entries / 500 estimated tokens injected. Root Operator can use `sortie_v010_reflection` to retain
336
+ verified process causes/preventions for later turns and sessions; this is not model training.
337
+ Storage and managed blocks are separate from stable. Set `reflection.enabled` to `false` to disable.
270
338
 
271
- The v0.10 host owns handoff and manifest controls under
339
+ The Mission host owns handoff and manifest controls under
272
340
  `.sortie-dogs-v010/contracts/`. Do not create a legacy root
273
341
  `operation-manifest.json` for this profile and do not edit generated controls.
274
342
  Delete `.sortie-dogs-v010/` only when no Sortie run is active.
275
343
 
276
344
  ### Validation policy
277
345
 
278
- `validationProfile` chooses non-canonical depth:
346
+ `validationProfile` chooses supplementary non-canonical depth:
279
347
 
280
348
  - `fast`: static checks
281
349
  - `balanced`: targeted checks
282
350
  - `assurance`: related checks
283
351
 
284
- Canonical proof remains canonical. Full-suite execution requires release context
285
- or explicit risk. Workers own static, targeted, and related checks; the root owns
286
- canonical and full-suite checks. An unchanged candidate, command, and environment
287
- reuse the same evidence instead of spending the validation budget again.
352
+ It does not replace meaningful declared formal checks or user/project-required broad validation.
353
+ Batch related edits, run focused checks, then execute every declared formal check in order on the
354
+ stable candidate. The implementing Worker or admitted correcting Reviewer runs those commands;
355
+ canonical/full-suite `owner=coordinator` is evidence accounting, not a requirement for root execution.
356
+
357
+ Keep required broad checks for the final integrated candidate. Reuse valid unchanged evidence only
358
+ when the contract permits; identity includes candidate, command, environment, scope and owner.
359
+ Required repeated occurrences retain their own execution identities and cannot be skipped as duplicates.
360
+ Native commands, actual working directory, exit, duration and saved source bindings establish freshness.
361
+ Diagnostics do not substitute for formal proof. Repeat checks when changes, failures or freshness require
362
+ it, rather than solely because a Worker changed or documentation was edited.
288
363
 
289
364
  ### Default routes
290
365
 
@@ -315,25 +390,40 @@ export { SortieDogsPlugin } from "sortie-dogs/plugin/stable";
315
390
 
316
391
  The stable profile uses `/sortie`, `dog-coordinator`,
317
392
  `.opencode/sortie-dogs.json`, `SORTIE_DOGS_CONFIG`, and `.sortie-dogs/`. Do not
318
- register stable and v0.10 from the same package installation path in one host.
393
+ register stable and `v010` from the same package installation path in one host.
319
394
 
320
395
  ## Global availability
321
396
 
322
- Project-local installation is recommended. To expose v0.10 assets globally:
397
+ Project-local installation is recommended. To expose the current Mission assets globally:
323
398
 
324
399
  ```sh
325
- npm install --global sortie-dogs
400
+ npm install --global sortie-dogs@0.13.4
326
401
  sortie-dogs init --global --profile v010
327
402
  ```
328
403
 
329
- Global initialization also merges `sortie-dogs` and `experimental.subagent_depth: 2` into the
330
- global OpenCode config, preserving unrelated settings and the default agent.
404
+ Global initialization registers the package or reuses an existing local V2 bridge, and sets subagent
405
+ depth to at least two, preserving unrelated settings and the default agent.
406
+
407
+ An existing `<OpenCode config root>/plugins/sortie-dogs/index.js` bridge importing `sortie-dogs/server`
408
+ can resolve a **separate dependency** under that config root. Updating npm-global alone does not update
409
+ it. For that layout, also install the same release at the actual config root, then rerun global init:
410
+
411
+ ```sh
412
+ npm install --prefix "$HOME/.config/opencode" sortie-dogs@0.13.4
413
+ sortie-dogs init --global --profile v010
414
+ ```
415
+
416
+ The command shows the default config root; use your actual root if overridden. For a configured npm
417
+ package entry, OpenCode V2 also provides `opencode plugin list` / `opencode plugin update`; exact
418
+ version pins require an explicit version change. Completely restart OpenCode after updating, then
419
+ check the installed package, asset marker and loaded plugin version.
331
420
 
332
421
  ## Updates and removal
333
422
 
334
- After replacing the dependency, rerun initialization and restart OpenCode:
423
+ For project-local updates, replace the dependency, rerun initialization and completely restart OpenCode:
335
424
 
336
425
  ```sh
426
+ npm install --save-dev sortie-dogs@latest
337
427
  npx sortie-dogs init .
338
428
  ```
339
429
 
@@ -341,6 +431,10 @@ npx sortie-dogs init .
341
431
  version, preserves user configuration, and stops safely on unknown ownership or
342
432
  conflicting files.
343
433
 
434
+ Align any exact version pin or separate bridge dependency with the intended release too. An installed
435
+ marker of `0.13.4-reviewer-continuous-v1` identifies the assets; it does not prove an already-running
436
+ OpenCode process has reloaded the plugin.
437
+
344
438
  There is no supported uninstall command. Remove the npm dependency separately,
345
439
  then follow the [safe manual removal guide](docs/uninstall.md). Delete only known
346
440
  Sortie-owned paths; never remove the whole `.opencode` directory or use broad
@@ -3,5 +3,5 @@
3
3
  * installed project marker without importing every asset body.
4
4
  */
5
5
  export declare const RUNTIME_ASSET_VERSION = "0.3.89-completion-proof-v1";
6
- export declare const V010_RUNTIME_ASSET_VERSION = "0.13.2-anko-recovery-v1";
6
+ export declare const V010_RUNTIME_ASSET_VERSION = "0.13.4-reviewer-continuous-v1";
7
7
  export type RuntimeAssetVersion = typeof RUNTIME_ASSET_VERSION | typeof V010_RUNTIME_ASSET_VERSION;
@@ -3,4 +3,4 @@
3
3
  * installed project marker without importing every asset body.
4
4
  */
5
5
  export const RUNTIME_ASSET_VERSION = "0.3.89-completion-proof-v1";
6
- export const V010_RUNTIME_ASSET_VERSION = "0.13.2-anko-recovery-v1";
6
+ export const V010_RUNTIME_ASSET_VERSION = "0.13.4-reviewer-continuous-v1";
@@ -6,3 +6,9 @@ export declare const CONTRACT_TEXT_LIMITS: Readonly<{
6
6
  command: 8192;
7
7
  path: 512;
8
8
  }>;
9
+ /** New Mission task generation only; persisted/legacy contracts retain their original bounds.
10
+ * Only target is authoring guidance. maximum is internal headroom, never a new target. */
11
+ export declare const MISSION_OBJECTIVE_LIMITS: Readonly<{
12
+ target: 2000;
13
+ maximum: 3000;
14
+ }>;
@@ -1,2 +1,5 @@
1
1
  /** Common text bounds shared by handoff, manifest and goal evidence validation. */
2
2
  export const CONTRACT_TEXT_LIMITS = Object.freeze({ title: 160, objective: 32768, statement: 1000, command: 8192, path: 512 });
3
+ /** New Mission task generation only; persisted/legacy contracts retain their original bounds.
4
+ * Only target is authoring guidance. maximum is internal headroom, never a new target. */
5
+ export const MISSION_OBJECTIVE_LIMITS = Object.freeze({ target: 2000, maximum: 3000 });
@@ -1,4 +1,4 @@
1
- import { type OperatorPlan, type OperatorState, type OperatorTask } from "./operator-runtime.js";
1
+ import { type OperatorPlan, type OperatorState, type OperatorTask, type OperatorRuntime } from "./operator-runtime.js";
2
2
  import { type RuntimeProfile } from "./runtime-profile.js";
3
3
  export declare const MISSION_REFERENCE = "SORTIE_MISSION_REF ";
4
4
  export declare const MISSION_REVIEW_REFERENCE = "SORTIE_MISSION_REVIEW_REF ";
@@ -44,7 +44,7 @@ export interface MissionAttempt {
44
44
  predecessorAttemptID?: string | null;
45
45
  /** Fingerprint of the settled scoped candidate against which a later Rescue is proposed. */
46
46
  candidateID?: string;
47
- kind: "implementation" | "normal_remediation" | "astra_rescue";
47
+ kind: "implementation" | "normal_remediation" | "astra_rescue" | "reviewer_correction" | "direct_execution";
48
48
  status: "pending" | "dispatched" | "succeeded" | "failed" | "cancelled" | "unconfirmed";
49
49
  callID?: string;
50
50
  childSessionID?: string;
@@ -100,6 +100,24 @@ export interface MissionLaunchConditions {
100
100
  applies_to: string;
101
101
  recordedAt?: string;
102
102
  }
103
+ /** Native terminal self-recheck by the correction author, explicitly not independent approval. */
104
+ export interface MissionSelfRecheck {
105
+ runID: string;
106
+ source: string;
107
+ /** Candidate source/check identity, excluding optional excerpt presentation. */
108
+ candidateSource?: string;
109
+ author: string;
110
+ callID: string;
111
+ promptID: string;
112
+ messageID: string;
113
+ nativeOutcome: "completed";
114
+ result: string;
115
+ unresolvedFindings: string[];
116
+ residualMajor?: {
117
+ reachable_path: string;
118
+ consequence: string;
119
+ };
120
+ }
103
121
  export interface OperatorMission {
104
122
  version: "0.12";
105
123
  id: string;
@@ -107,6 +125,12 @@ export interface OperatorMission {
107
125
  requests: MissionRequest[];
108
126
  /** Prior public conversation context, not additional immutable requirements. */
109
127
  context?: MissionContext[];
128
+ /** Explicit continue deliveries, keyed by the original real turn. */
129
+ steering?: {
130
+ requestID: string;
131
+ child: string;
132
+ status: "pending" | "queued";
133
+ }[];
110
134
  kind?: "implementation" | "operation";
111
135
  /** Native shell observations of the requested operation, separate from auxiliary checks. */
112
136
  execution?: MissionExecution;
@@ -127,6 +151,15 @@ export interface OperatorMission {
127
151
  runID: string | null;
128
152
  /** Git HEAD before this mission's first implementation unit, retained across replans and commits. */
129
153
  reviewBaseline?: string;
154
+ /** A dated host Git observation, never an inference from a Worker report or lifecycle plan. */
155
+ deliveryObservation?: {
156
+ run_id: string;
157
+ observed_at: string;
158
+ head: string | null;
159
+ branch: string | null;
160
+ clean: boolean;
161
+ source: string;
162
+ };
130
163
  /** Cumulative declared review inputs/outputs, including units that failed after writing source. */
131
164
  reviewScope?: MissionReviewScope;
132
165
  /** Optional Advisor/Scout decisions and native consultation outcomes; never an admission gate. */
@@ -152,17 +185,43 @@ export interface OperatorMission {
152
185
  runID: string;
153
186
  risk: string[];
154
187
  source: string;
188
+ candidateSource?: string;
155
189
  task: OperatorTask | null;
190
+ callID?: string;
156
191
  evidence?: MissionEvidenceExcerpt[];
157
192
  requestFingerprint?: string;
158
- verdict: "pending" | "PASS" | "findings" | "evidence-gaps" | "skipped-low-risk";
193
+ verdict: "pending" | "PASS" | "findings" | "evidence-gaps" | "skipped-low-risk" | "self-rechecked";
159
194
  result?: string;
160
195
  child?: string;
196
+ mode?: "independent" | "self-recheck";
197
+ admittedAt?: number;
198
+ promptID?: string;
199
+ selfRecheck?: MissionSelfRecheck;
161
200
  /** Completed, independent initial review for this mission, not merely an inherited child ID. */
162
201
  initialPrompt?: string;
163
202
  /** Observed evidence-only reviews; reporting only, never an acceptance threshold. */
164
203
  evidenceGapReviews?: number;
165
204
  };
205
+ /** Retained independently of later review generations; authors never become independent reviewers. */
206
+ corrections?: {
207
+ author: string;
208
+ reviewIdentity: string;
209
+ priorRunID: string;
210
+ runID: string;
211
+ priorSource: string;
212
+ findings: string;
213
+ initialPrompt: string;
214
+ baseline?: string;
215
+ /** Actual still-running initial Review; direct correction is not another child terminal. */
216
+ inlineReview?: {
217
+ callID: string;
218
+ promptID: string;
219
+ admittedAt: number;
220
+ reviewIdentity: string;
221
+ };
222
+ status: "prepared" | "running" | "ready" | "failed" | "cancelled";
223
+ selfRecheck?: MissionSelfRecheck;
224
+ }[];
166
225
  }
167
226
  export declare const MISSION_CONSULTATION_LIMIT = 32;
168
227
  /** Review coverage survives a narrower replan; it is not a Worker write grant. */
@@ -171,6 +230,8 @@ export declare function missionReviewScope(previous: MissionReviewScope | undefi
171
230
  export declare function missionReviewVerdict(text: string): "PASS" | "evidence-gaps" | "findings";
172
231
  /** Whether the recorded review permits submission and acceptance of the current candidate. */
173
232
  export declare function missionReviewAccepted(review: NonNullable<OperatorMission["review"]>): boolean;
233
+ /** A short native report, not a tag/hash-based second-review policy or an approval checklist. */
234
+ export declare function missionSelfRecheckReport(text: string, source: string, hostBound?: boolean): Pick<MissionSelfRecheck, "unresolvedFindings" | "residualMajor"> | undefined;
174
235
  export declare function missionExecutionStatus(mission: OperatorMission): "not-required" | "not-started" | "running" | "execution-failed" | "executed";
175
236
  /** Observe the command's own terminal summary, never Worker/Reviewer prose. Scores are result data,
176
237
  * not process success. Unstructured commands retain their native exit semantics. */
@@ -187,6 +248,7 @@ export declare class OperatorMissionRuntime {
187
248
  private readonly writes;
188
249
  constructor(projectRoot: string, profile: RuntimeProfile);
189
250
  private file;
251
+ correctionReference(root: string, reviewIdentity: string): string;
190
252
  private load;
191
253
  /** Recover non-replacement continuations only; an explicit replacement must link the current cancelled run. */
192
254
  private loadMission;
@@ -228,7 +290,10 @@ export declare class OperatorMissionRuntime {
228
290
  brief(state: OperatorMission): string;
229
291
  }
230
292
  /** The model supplies only useful unit facts; IDs, proof projection and control documents are generated here. */
231
- export declare function missionPlan(mission: OperatorMission, raw: unknown): OperatorPlan;
293
+ export declare function missionPlan(mission: OperatorMission, raw: unknown, projectRoot?: string): OperatorPlan;
294
+ /** Compact provenance, not a new acceptance verdict or a substitute for original-request comparison. */
295
+ export declare function missionAcceptanceSummary(mission: OperatorMission, run: OperatorState | undefined, operators: OperatorRuntime, observe?: (validation: readonly string[], child: string | null, notBefore?: number) => Promise<unknown>): Promise<Record<string, unknown>>;
296
+ export declare function missionReviewIndependent(mission: OperatorMission, child: string | undefined): boolean;
232
297
  export declare function missionPacket(mission: OperatorMission, run?: OperatorState): Record<string, unknown>;
233
298
  /** Models forward a short capability, never recopy the host's evidence hashes and source packet. */
234
299
  export declare function missionReviewTask(mission: OperatorMission): OperatorTask;