sortie-dogs 0.13.3 → 0.13.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -35,6 +35,11 @@ implementation, validation, review, and model routing.
35
35
  Guides: [日本語](docs/guide-ja.md) · [简体中文](docs/guide-zh-CN.md) ·
36
36
  [Testing](docs/testing.md) · [CLI testing](docs/cli-testing.md)
37
37
 
38
+ **Current release: [v0.13.5](https://github.com/zufall-upon/Sortie-dogs/releases/tag/v0.13.5)**
39
+ ([release notes](docs/release-v0.13.5.md)). The default Mission runtime retains the `v010`
40
+ profile, command and configuration names for compatibility; these names do not mean v0.10 is installed.
41
+ The current asset marker is `0.13.5-live-state-v1`.
42
+
38
43
  ## SWE-bench Lite: 170/300 (56.67%)
39
44
 
40
45
  The fixed **Sortie-dogs v0.12.24** harness resolved **170 of 300 SWE-bench Lite test issues** in one pass@1 campaign, with 9 empty patches and no official evaluation errors. Every instance has a frozen prediction and an inference-time trajectory. The task Workers ran `openai/gpt-6-luna-fast#max`; operator, coordinator and review roles ran `openai/gpt-6-sol#xhigh`. This is a system result, **not** a Luna-only model comparison or a Verified/full SWE-bench score.
@@ -43,22 +48,25 @@ The fixed **Sortie-dogs v0.12.24** harness resolved **170 of 300 SWE-bench Lite
43
48
 
44
49
  The single official 300-instance report and frozen predictions are hash-bound in the report. Confirmed inference expense was **$162.99**; a separate **$34.60** of usage has unknown pricing and is held against the campaign cap, **not** counted as known expense. Leaderboard registration and maintainer acceptance are separate from this official local evaluation.
45
50
 
46
- > **Beta:** v0.12.2 builds on the v0.10.23 execution engine. Runtime behavior,
51
+ Historical scores below belong to their fixed candidates, not v0.13.5. SWE-bench is a separate,
52
+ optional measurement rather than a mandatory release gate.
53
+
54
+ > **Beta:** v0.13.x is still stabilizing. Runtime behavior,
47
55
  > configuration, and generated assets may still change before 1.0.
48
56
 
49
57
  ## Quick start
50
58
 
51
- Requirements: Node.js 22.6 or newer, npm, and OpenCode.
59
+ Requirements: Node.js 22.6 or newer, npm, and OpenCode V2.
52
60
 
53
61
  Run these commands in the target project:
54
62
 
55
63
  ```sh
56
- npm install --save-dev sortie-dogs
64
+ npm install --save-dev sortie-dogs@latest
57
65
  npx sortie-dogs init .
58
66
  ```
59
67
 
60
- The package retains the `v010` profile/namespace for compatibility. `init` adds the OpenCode V2 plugin and the required
61
- two-level subagent depth to `.opencode/opencode.json(c)`, preserving existing settings:
68
+ `init` defaults to the `v010` Mission profile and registers the OpenCode V2 plugin in
69
+ `.opencode/opencode.json(c)`, preserving unrelated settings. It sets subagent depth to at least two:
62
70
 
63
71
  ```json
64
72
  {
@@ -67,51 +75,88 @@ two-level subagent depth to `.opencode/opencode.json(c)`, preserving existing se
67
75
  }
68
76
  ```
69
77
 
70
- Restart OpenCode, then run:
78
+ If an existing local bridge already imports `sortie-dogs/server`, `init` reuses it instead of adding
79
+ a duplicate package entry. A larger existing subagent depth is retained.
80
+
81
+ Completely restart OpenCode, then run:
71
82
 
72
83
  ```text
73
84
  /sortie-v010 <task>
74
85
  ```
75
86
 
76
87
  Selecting `dog-operator` directly starts the same workflow. `dog-operator` is the
77
- only user-facing v0.10 authority. `dogs-coordinator` and every `*-v010` role are
88
+ user-facing entry point for the Mission profile. `dogs-coordinator` and every `*-v010` role are
78
89
  internal children and must not be selected as task entry points.
79
90
 
80
- `init` installs runtime assets and merges the required OpenCode settings; the `plugins` entry loads runtime enforcement and
81
- model routing. A new session alone does not reload an updated
82
- plugin process, so restart OpenCode after installation or upgrade.
83
-
84
- ## v0.13.3 runtime updates
85
-
86
- The current release integrates native background Mission execution, concise authoritative Worker
87
- handoffs, and same-context Reviewer corrections from PRs #148/#149 and their Ubuntu remediation.
88
- After finding defects, the original Reviewer can correct and explicitly self-recheck in the same
89
- native session, retaining formal checks, current-source evidence, Git delivery and cumulative budget.
90
- Author self-recheck is recorded as non-independent; a different Reviewer is conditional on a concrete
91
- reachable residual Major risk. Unresolved Major or Medium findings still block acceptance.
92
-
93
- The installed native correction fixture succeeded, but the original Anko measurements remain
94
- unaccepted, including the latest 25-minute run. This release does not claim general speedup,
95
- original-task completion or a new SWE-bench score. See the [release notes](docs/release-v0.13.3.md)
96
- and [retained implementation and measurement history](docs/anko-pr149-ubuntu-handoff.md).
97
-
98
- ## v0.12.0 workflow
99
-
100
- v0.12.0 keeps Operator → Coordinator → Worker, with independent review:
91
+ `init` installs runtime assets and merges the required OpenCode settings; the package entry or
92
+ existing local bridge loads enforcement and model routing. OpenCode can reload watched configuration,
93
+ but replacing an installed dependency may require a full restart. A new chat session alone does not
94
+ prove the newly installed plugin is loaded.
95
+
96
+ ## v0.13.5 runtime updates
97
+
98
+ PR #154 preserves stable model instructions and appends changed host state at native history
99
+ boundaries, with complete current-state reconstruction after compaction. Genuine Reviewer tools
100
+ have deterministic read-only-first ordering; the real correction-permission transition still remains.
101
+ The final candidate restores v3 behavioral review and combined assignment/findings while retaining
102
+ exact instruction discovery, inherited Reviewer formal-command delivery, literal local shell-file
103
+ scope reconciliation and explicit repository-root read/write scope support. New whole-project
104
+ captures avoid bookkeeping-only invalidation; legacy evidence and real source/artifact freshness remain.
105
+
106
+ The final pre-release cycle 11 Anko sample passed official local score 1 and selected public probes 9/9
107
+ in 26m1.866s at estimated $1.41968316. It was faster but 15.762% costlier than the earlier v3 sample,
108
+ not combined cost-preserving optimization or a general quality/speed guarantee. Release receipts:
109
+ `_testenv/releases/0.13.5/`. See [release notes](docs/release-v0.13.5.md) and
110
+ [candidate tradeoffs/failures](docs/cache-prefix-loop-20261003.md). Native Worker startup is not completion.
111
+
112
+ ## v0.13.4 runtime updates (retained)
113
+
114
+ PR #152 restores same-session Coordinator/Operator implementation and formal validation through
115
+ `plan_units(executor="self")`, `start_direct_unit` and `finish_direct_unit`. Known single-unit work can
116
+ combine Mission start and planning; Luna Fast/max Worker routing remains the default. Full generated
117
+ contracts remain visible when an explicit Read line range covers the file. Native background
118
+ responsiveness remains: a launch acknowledgement or idle root is not Mission completion.
119
+
120
+ The initial independent Reviewer can investigate, correct, formally validate, deliver and self-recheck
121
+ continuously in its original native Task. Later Major/Medium findings accumulate in that same correction
122
+ context. Author self-recheck remains `self-rechecked`, `independent=false`, never independent `PASS`.
123
+ A different Reviewer is conditional on concrete residual Major risk; unresolved Major/Medium findings
124
+ still block acceptance. Operator owns final comparison and receipt.
125
+
126
+ Inherited compiler scratch no longer falsely invalidates broad-scope formal proof. Host-observed Git
127
+ delivery, caller-setting review and test-composition guidance reduce avoidable detours. Saved host
128
+ completion cards remain in tool history/UI, while outgoing V2 model/compaction context omits only their
129
+ presentation body, retaining receipt and evidence identities.
130
+
131
+ Two runs of the same fixed pre-release v8 package completed the original Anko task with official local
132
+ score 1 (F2P 9/9, P2P 94/94) and selected public probes 9/9. Times were 26m49s and 25m55s, costs
133
+ $1.47908048 and $1.54610816; each saved more than seven minutes versus the recorded v5 sample.
134
+ These limited same-task observations do not establish general speedup or a new SWE-bench score.
135
+ Release preflight, full tests, fixed-commit Windows CI and native Worker-start receipts are retained in
136
+ `_testenv/releases/0.13.4/`; startup/model identity is not task completion. See the
137
+ [release notes](docs/release-v0.13.4.md) and [quality-loop evidence](docs/nightly-quality-loop-20261002.md).
138
+
139
+ ## Mission workflow
140
+
141
+ Use Operator → Worker when one useful unit and its meaningful formal check are known;
142
+ use Operator → Coordinator → Worker for actual discovery or decomposition:
101
143
 
102
144
  - `dog-operator` states a few requirements/negative constraints and owns user decisions and final acceptance.
103
145
  The host saves the original user message verbatim.
104
146
  - Hidden `dogs-coordinator` owns investigation, unit declarations, Worker/Scout/Advisor/Reviewer dispatch,
105
147
  in-request write-scope extensions, and corrections. It can read/search and run confirmation shell commands;
106
- source editing tools belong to Worker.
148
+ it can implement and formally validate directly in its own session, or delegate a unit to Worker.
107
149
  - `dog-worker-v010` implements a host-generated unit within its file/directory write scopes.
108
150
  Investigation commands need no pre-registration; formal checks retain real host-recorded results.
109
- - High-risk changes require an independent Reviewer. Low-risk skips are explicit and recorded.
110
- - Simple low-risk single-unit work retains Operator → Worker Fast-lane dispatch.
151
+ - High-risk changes require an initial independent Reviewer with read/search access. Low-risk skips
152
+ are explicit and recorded. Reviewer-owned corrections follow the self-recheck policy above.
153
+ - Fast-lane can include high-risk single-unit work; it never implies a review skip.
154
+ - Investigation, edits, formal checks and requested Git delivery stay in the same implementing child.
155
+ Explicit user ordering is retained; no routine plan-approval or commit-only handoff is needed.
111
156
  - Unit progress appears on the running Task without stopping Coordinator or prompting Operator.
112
157
 
113
- The v0.10 profile is serial by design. The stable profile's Luna fabric and
114
- parallel integration path are not exposed in this profile. More agents are not a
158
+ The `v010` Mission profile is serial by design; background responsiveness does not add parallel writers.
159
+ The stable profile's Luna fabric and parallel integration path are not exposed in this profile. More agents are not a
115
160
  goal; preserving quality while reducing unnecessary expensive work is.
116
161
 
117
162
  ### SWE-bench evaluation
@@ -186,32 +231,51 @@ astroid **3/5**, pyvista **0/1**, sqlfluff **1/5**.
186
231
 
187
232
  ## Mission tools
188
233
 
189
- 1. `start_mission`: Operator supplies concise requirements; the host saves original messages and returns a Coordinator task.
190
- 2. `plan_units`: Coordinator supplies title, objective, file/directory scopes and formal checks. The host generates
191
- IDs, handoff, manifest, proof mapping and the ready Worker task. No proposal approval round trip.
192
- 3. `operator_next`: advance serial units. `expand_unit` or a reasoned `plan_units` correction extends/replaces
193
- settled execution within the original requirements and retained cumulative budget.
234
+ 1. `start_mission`: Operator supplies concise requirements; the host saves original messages. A known single unit
235
+ can include `unit` to combine start/planning and return its configured Worker task.
236
+ 2. `plan_units`: Operator or Coordinator supplies title, objective, file/directory scopes and formal checks. The host generates
237
+ IDs, handoff, manifest, proof mapping and the ready Worker task. `executor="self"` keeps execution in the
238
+ same controller session; `start_direct_unit` / `finish_direct_unit` retain observed formal-check freshness.
239
+ 3. `operator_next`: advance serial units. `expand_unit` reconciles required in-request outputs while
240
+ preserving the same Task. A reasoned `plan_units` correction or `retry_mission_unit` handles ordinary
241
+ unit recovery under the original requirements and cumulative budget.
194
242
  4. `review_mission`: generate the independent review packet from source, requirements and observed checks;
195
243
  dispatch its Reviewer task for high-risk changes or record a low-risk skip.
196
- 5. `submit_mission`: Coordinator returns a completion candidate, user-only decision, or proven external/scope/budget blocker.
197
- 6. `complete_mission`: Operator compares the original request, source and evidence, then explicitly accepts.
244
+ 5. `repair_review`: record findings and continue correction in the running original Reviewer Task, or resume
245
+ that same native session. Later findings accumulate; run inherited checks/requested Git delivery and
246
+ explicitly report `SELF_RECHECKED` in that same Task. Legacy
247
+ `CORRECTION_READY` alone requires a same-author read-only fallback through `review_mission`.
248
+ 6. `submit_mission`: Coordinator returns a completion candidate, user-only decision, or proven external/scope/budget blocker.
249
+ 7. `complete_mission`: Operator compares the original request, source and evidence, then explicitly accepts.
198
250
  Only a succeeded receipt authorizes DONE and the measured 🐾 return report.
199
251
 
200
252
  All tool names use the `sortie_v010_` prefix. Prior proposal/plan-repair tools remain in the compatibility
201
- implementation but are hidden from the normal v0.12 model tool list. An in-request path extension is a
202
- Coordinator decision; changing the original requirements or increasing budget returns to Operator/user.
253
+ implementation but are hidden from the normal Mission tool list. Operator/Coordinator handle in-request
254
+ path reconciliation without a new user approval. Changes beyond the original requirements or cumulative
255
+ budget return to Operator/user. `EVIDENCE_GAPS` is an advisory limitation, not Review `PASS` or an
256
+ automatic extra review; failed or missing required checks still prevent acceptance.
203
257
 
204
258
  Durable profile state and hash-bound task references support restart and
205
259
  compaction recovery without reconstructing criteria from summary prose. Stale,
206
- foreign-root, or changed references are rejected. An optional Git lifecycle can
207
- create one non-overwriting branch and one explicit-path commit; it never grants
208
- arbitrary Git, force push, release, or publication authority.
260
+ foreign-root, or changed references are rejected. Requested `git add <paths>` and `git commit -m ...`
261
+ use the actual source write scope, not a fabricated `.git/**` scope. An optional host-managed Git
262
+ lifecycle also retains its branch, commit and post-commit boundaries. Neither mode grants arbitrary
263
+ Git, force push, release or publication authority.
264
+
265
+ ### Progress and acceptance evidence
266
+
267
+ `sortie_v010_operator_status` keeps original requests, formal command/exit/timing observations,
268
+ review disposition and recorded delivery in a compact Mission view. `{ "view": "progress" }` exposes
269
+ the current unit, completed/total units, host budget and next action; `{ "view": "full" }` or
270
+ `details_ref` provides full snapshot diagnostics. Unknown clean state or a failed commit is not delivery
271
+ success. Reading progress does not dispatch, retry or accept work; use native completion notifications
272
+ instead of polling. Existing status reconciliation can recover a missed child terminal event.
209
273
 
210
274
  ## Configuration
211
275
 
212
276
  ### Profile files and precedence
213
277
 
214
- The default package entry is the v0.10 profile:
278
+ The default package entry is the `v010` Mission profile:
215
279
 
216
280
  - Command: `/sortie-v010`
217
281
  - Primary agent: `dog-operator`
@@ -221,9 +285,11 @@ The default package entry is the v0.10 profile:
221
285
  - Runtime state: `.sortie-dogs-v010/`
222
286
  - Installed asset marker: `.opencode/sortie-dogs-v010.version`
223
287
 
288
+ For a global install, the marker is `<OpenCode config root>/sortie-dogs-v010.version`.
289
+
224
290
  Precedence is built-in defaults, global file, project file, environment JSON,
225
291
  then plugin factory options. Unknown properties or invalid types are rejected.
226
- Use external v0.10 role names such as `dog-operator`, `dogs-coordinator`, and
292
+ Use external `v010` role names such as `dog-operator`, `dogs-coordinator`, and
227
293
  `dog-reviewer-v010` in `modelRouting`; do not also declare their stable aliases.
228
294
 
229
295
  Example `.opencode/sortie-dogs-v010.json`:
@@ -266,8 +332,8 @@ not invent, probe, or translate variant names.
266
332
  - `freeTierFallbackModels`: ordered global last-resort model IDs. Default:
267
333
  `opencode/deepseek-v4-flash-free`; `[]` disables this fallback.
268
334
  - `dedicatedWorkerModel`: canonical stable serial target, default
269
- `openai/gpt-6.1-sol` / `medium`. The v0.10 profile also supplies its explicit
270
- role routes below; do not infer the v0.10 worker route from this stable setting.
335
+ `openai/gpt-6.1-sol` / `medium`. The Mission profile supplies its explicit
336
+ role routes below; do not infer its Worker route from this stable setting.
271
337
  - `consultation.strategy`: fixed advisor identity, optional `required`, and
272
338
  positive `maxCallsPerCandidate`; default one call and not required.
273
339
  - `consultation.sourceReview`: risk-based review with `maxCallsPerCandidate`
@@ -281,27 +347,35 @@ not invent, probe, or translate variant names.
281
347
  - `continuation.summarizeModel`: optional explicit compaction model; omission
282
348
  reuses the latest observed root model.
283
349
  - `validationProfile`: `fast`, `balanced`, or `assurance`; default `balanced`.
284
- - `reflection`: accepted by the shared schema, but reflection writes are not
285
- exposed by the serial v0.10 profile. Stable reflection remains opt-in and off by
286
- default.
350
+ - `reflection`: enabled by default for `run`, `project` and `global` layers, with at most three
351
+ entries / 500 estimated tokens injected. Root Operator can use `sortie_v010_reflection` to retain
352
+ verified process causes/preventions for later turns and sessions; this is not model training.
353
+ Storage and managed blocks are separate from stable. Set `reflection.enabled` to `false` to disable.
287
354
 
288
- The v0.10 host owns handoff and manifest controls under
355
+ The Mission host owns handoff and manifest controls under
289
356
  `.sortie-dogs-v010/contracts/`. Do not create a legacy root
290
357
  `operation-manifest.json` for this profile and do not edit generated controls.
291
358
  Delete `.sortie-dogs-v010/` only when no Sortie run is active.
292
359
 
293
360
  ### Validation policy
294
361
 
295
- `validationProfile` chooses non-canonical depth:
362
+ `validationProfile` chooses supplementary non-canonical depth:
296
363
 
297
364
  - `fast`: static checks
298
365
  - `balanced`: targeted checks
299
366
  - `assurance`: related checks
300
367
 
301
- Canonical proof remains canonical. Full-suite execution requires release context
302
- or explicit risk. Workers own static, targeted, and related checks; the root owns
303
- canonical and full-suite checks. An unchanged candidate, command, and environment
304
- reuse the same evidence instead of spending the validation budget again.
368
+ It does not replace meaningful declared formal checks or user/project-required broad validation.
369
+ Batch related edits, run focused checks, then execute every declared formal check in order on the
370
+ stable candidate. The implementing Worker or admitted correcting Reviewer runs those commands;
371
+ canonical/full-suite `owner=coordinator` is evidence accounting, not a requirement for root execution.
372
+
373
+ Keep required broad checks for the final integrated candidate. Reuse valid unchanged evidence only
374
+ when the contract permits; identity includes candidate, command, environment, scope and owner.
375
+ Required repeated occurrences retain their own execution identities and cannot be skipped as duplicates.
376
+ Native commands, actual working directory, exit, duration and saved source bindings establish freshness.
377
+ Diagnostics do not substitute for formal proof. Repeat checks when changes, failures or freshness require
378
+ it, rather than solely because a Worker changed or documentation was edited.
305
379
 
306
380
  ### Default routes
307
381
 
@@ -332,25 +406,40 @@ export { SortieDogsPlugin } from "sortie-dogs/plugin/stable";
332
406
 
333
407
  The stable profile uses `/sortie`, `dog-coordinator`,
334
408
  `.opencode/sortie-dogs.json`, `SORTIE_DOGS_CONFIG`, and `.sortie-dogs/`. Do not
335
- register stable and v0.10 from the same package installation path in one host.
409
+ register stable and `v010` from the same package installation path in one host.
336
410
 
337
411
  ## Global availability
338
412
 
339
- Project-local installation is recommended. To expose v0.10 assets globally:
413
+ Project-local installation is recommended. To expose the current Mission assets globally:
340
414
 
341
415
  ```sh
342
- npm install --global sortie-dogs
416
+ npm install --global sortie-dogs@0.13.5
343
417
  sortie-dogs init --global --profile v010
344
418
  ```
345
419
 
346
- Global initialization also merges `sortie-dogs` and `experimental.subagent_depth: 2` into the
347
- global OpenCode config, preserving unrelated settings and the default agent.
420
+ Global initialization registers the package or reuses an existing local V2 bridge, and sets subagent
421
+ depth to at least two, preserving unrelated settings and the default agent.
422
+
423
+ An existing `<OpenCode config root>/plugins/sortie-dogs/index.js` bridge importing `sortie-dogs/server`
424
+ can resolve a **separate dependency** under that config root. Updating npm-global alone does not update
425
+ it. For that layout, also install the same release at the actual config root, then rerun global init:
426
+
427
+ ```sh
428
+ npm install --prefix "$HOME/.config/opencode" sortie-dogs@0.13.5
429
+ sortie-dogs init --global --profile v010
430
+ ```
431
+
432
+ The command shows the default config root; use your actual root if overridden. For a configured npm
433
+ package entry, OpenCode V2 also provides `opencode plugin list` / `opencode plugin update`; exact
434
+ version pins require an explicit version change. Completely restart OpenCode after updating, then
435
+ check the installed package, asset marker and loaded plugin version.
348
436
 
349
437
  ## Updates and removal
350
438
 
351
- After replacing the dependency, rerun initialization and restart OpenCode:
439
+ For project-local updates, replace the dependency, rerun initialization and completely restart OpenCode:
352
440
 
353
441
  ```sh
442
+ npm install --save-dev sortie-dogs@latest
354
443
  npx sortie-dogs init .
355
444
  ```
356
445
 
@@ -358,6 +447,10 @@ npx sortie-dogs init .
358
447
  version, preserves user configuration, and stops safely on unknown ownership or
359
448
  conflicting files.
360
449
 
450
+ Align any exact version pin or separate bridge dependency with the intended release too. An installed
451
+ marker of `0.13.5-live-state-v1` identifies the assets; it does not prove an already-running
452
+ OpenCode process has reloaded the plugin.
453
+
361
454
  There is no supported uninstall command. Remove the npm dependency separately,
362
455
  then follow the [safe manual removal guide](docs/uninstall.md). Delete only known
363
456
  Sortie-owned paths; never remove the whole `.opencode` directory or use broad
@@ -3,5 +3,5 @@
3
3
  * installed project marker without importing every asset body.
4
4
  */
5
5
  export declare const RUNTIME_ASSET_VERSION = "0.3.89-completion-proof-v1";
6
- export declare const V010_RUNTIME_ASSET_VERSION = "0.13.3-reviewer-context-v1";
6
+ export declare const V010_RUNTIME_ASSET_VERSION = "0.13.5-live-state-v1";
7
7
  export type RuntimeAssetVersion = typeof RUNTIME_ASSET_VERSION | typeof V010_RUNTIME_ASSET_VERSION;
@@ -3,4 +3,4 @@
3
3
  * installed project marker without importing every asset body.
4
4
  */
5
5
  export const RUNTIME_ASSET_VERSION = "0.3.89-completion-proof-v1";
6
- export const V010_RUNTIME_ASSET_VERSION = "0.13.3-reviewer-context-v1";
6
+ export const V010_RUNTIME_ASSET_VERSION = "0.13.5-live-state-v1";
@@ -6,7 +6,8 @@ export declare const CONTRACT_TEXT_LIMITS: Readonly<{
6
6
  command: 8192;
7
7
  path: 512;
8
8
  }>;
9
- /** New Mission task generation only; persisted/legacy contracts retain their original bounds. */
9
+ /** New Mission task generation only; persisted/legacy contracts retain their original bounds.
10
+ * Only target is authoring guidance. maximum is internal headroom, never a new target. */
10
11
  export declare const MISSION_OBJECTIVE_LIMITS: Readonly<{
11
12
  target: 2000;
12
13
  maximum: 3000;
@@ -1,4 +1,5 @@
1
1
  /** Common text bounds shared by handoff, manifest and goal evidence validation. */
2
2
  export const CONTRACT_TEXT_LIMITS = Object.freeze({ title: 160, objective: 32768, statement: 1000, command: 8192, path: 512 });
3
- /** New Mission task generation only; persisted/legacy contracts retain their original bounds. */
3
+ /** New Mission task generation only; persisted/legacy contracts retain their original bounds.
4
+ * Only target is authoring guidance. maximum is internal headroom, never a new target. */
4
5
  export const MISSION_OBJECTIVE_LIMITS = Object.freeze({ target: 2000, maximum: 3000 });
@@ -39,6 +39,9 @@ export interface GoalEvidence {
39
39
  readonly candidate_paths: readonly string[];
40
40
  /** Missing on legacy evidence: retain its original all-paths snapshot recipe. */
41
41
  readonly source_policy?: "project-files-v1" | "declared-paths-v1";
42
+ /** New captures distinguish a whole-project output grant from explicitly named control-like artifacts.
43
+ * Missing on old evidence: keep its original candidate recipe, including bookkeeping bytes. */
44
+ readonly candidate_policy?: "project-root-artifacts-v1";
42
45
  /** Fixed when validation starts. Full manifest_hash still identifies the historical execution contract. */
43
46
  readonly freshness?: {
44
47
  readonly contract_hash: string;
@@ -65,6 +65,7 @@ export function validGoalEvidence(value, state) {
65
65
  protectedBinding.candidate_paths.every(text) &&
66
66
  (protectedBinding.source_policy === undefined || protectedBinding.source_policy === "project-files-v1" ||
67
67
  protectedBinding.source_policy === "declared-paths-v1") &&
68
+ (protectedBinding.candidate_policy === undefined || protectedBinding.candidate_policy === "project-root-artifacts-v1") &&
68
69
  (protectedBinding.freshness === undefined || (protectedBinding.freshness !== null && typeof protectedBinding.freshness === "object" &&
69
70
  HASH.test(protectedBinding.freshness.contract_hash) &&
70
71
  Array.isArray(protectedBinding.freshness.scratch_paths) && protectedBinding.freshness.scratch_paths.every(text) &&
@@ -44,7 +44,7 @@ export interface MissionAttempt {
44
44
  predecessorAttemptID?: string | null;
45
45
  /** Fingerprint of the settled scoped candidate against which a later Rescue is proposed. */
46
46
  candidateID?: string;
47
- kind: "implementation" | "normal_remediation" | "astra_rescue" | "reviewer_correction";
47
+ kind: "implementation" | "normal_remediation" | "astra_rescue" | "reviewer_correction" | "direct_execution";
48
48
  status: "pending" | "dispatched" | "succeeded" | "failed" | "cancelled" | "unconfirmed";
49
49
  callID?: string;
50
50
  childSessionID?: string;
@@ -151,6 +151,15 @@ export interface OperatorMission {
151
151
  runID: string | null;
152
152
  /** Git HEAD before this mission's first implementation unit, retained across replans and commits. */
153
153
  reviewBaseline?: string;
154
+ /** A dated host Git observation, never an inference from a Worker report or lifecycle plan. */
155
+ deliveryObservation?: {
156
+ run_id: string;
157
+ observed_at: string;
158
+ head: string | null;
159
+ branch: string | null;
160
+ clean: boolean;
161
+ source: string;
162
+ };
154
163
  /** Cumulative declared review inputs/outputs, including units that failed after writing source. */
155
164
  reviewScope?: MissionReviewScope;
156
165
  /** Optional Advisor/Scout decisions and native consultation outcomes; never an admission gate. */
@@ -203,6 +212,13 @@ export interface OperatorMission {
203
212
  findings: string;
204
213
  initialPrompt: string;
205
214
  baseline?: string;
215
+ /** Actual still-running initial Review; direct correction is not another child terminal. */
216
+ inlineReview?: {
217
+ callID: string;
218
+ promptID: string;
219
+ admittedAt: number;
220
+ reviewIdentity: string;
221
+ };
206
222
  status: "prepared" | "running" | "ready" | "failed" | "cancelled";
207
223
  selfRecheck?: MissionSelfRecheck;
208
224
  }[];
@@ -413,8 +413,8 @@ export class OperatorMissionRuntime {
413
413
  }
414
414
  brief(state) {
415
415
  return [`mission_id: ${state.id}`, `project_root: ${this.projectRoot}`, "Use the user's language below for all replies and Task titles.",
416
- "Own investigation, unit declarations, Worker/Scout/Advisor/independent Reviewer dispatch, scope extensions and corrections.",
417
- "Start the first useful Worker promptly. No proposal/approval phase. Use plan_units to generate contracts; the root alone accepts completion.",
416
+ "Own investigation, implementation, formal validation, corrections and Worker/Scout/Advisor/independent Reviewer dispatch.",
417
+ "Use plan_units with executor=self for direct work in this session, or delegate promptly when useful. finish_direct_unit records native checks without a Worker handoff. No proposal/approval phase; root alone accepts completion.",
418
418
  "Escalate only a completion candidate, a user-only decision, or an extension of original requirements/budget. Unit progress is published without stopping you.",
419
419
  "Requirements:", ...state.requirements.map(item => `${item.id}: ${item.text}`),
420
420
  `Confirmed launch conditions (fixed limits, not consumption or remaining budget): ${JSON.stringify(state.launchConditions ?? [])}`,
@@ -443,10 +443,11 @@ export function missionPlan(mission, raw, projectRoot) {
443
443
  if (!Array.isArray(entries) || !entries.every(item => typeof item === "string"))
444
444
  throw new Error(`mission-unit-${index + 1}: ${field} must be paths`);
445
445
  return [...new Set(entries.map(item => {
446
- // A model's repository-root read means the current project, not an invalid empty path.
447
- // Resolve only this read shorthand at the host boundary; saved plans and write scopes stay exact.
448
- const rootRead = field === "read" && [".", "./", ".\\", "./**", ".\\**"].includes(item);
449
- return normalizeExecutionScope(rootRead && projectRoot ? `${resolve(projectRoot).replaceAll("\\", "/")}/**` : item);
446
+ // An explicit repository-root scope means this project for both reads and writes.
447
+ // Resolve this shorthand only at the host boundary; saved plans, prohibitions and
448
+ // native permissions still use the exact directory scope, never an inferred grant.
449
+ const rootScope = [".", "./", ".\\", "./**", ".\\**"].includes(item);
450
+ return normalizeExecutionScope(rootScope && projectRoot ? `${resolve(projectRoot).replaceAll("\\", "/")}/**` : item);
450
451
  }))];
451
452
  };
452
453
  // A sole unit owns the whole request. This schedules work; it does not prove acceptance.
@@ -508,7 +509,9 @@ export async function missionAcceptanceSummary(mission, run, operators, observe)
508
509
  return [];
509
510
  return (unit.evidence ?? []).map(proof => ({ run_id: state.runID, unit_id: unit.unit.id,
510
511
  task_id: unit.task.prompt.match(/^task_id: (.+)$/mu)?.[1] ?? null,
511
- worker_session_id: unit.childSessionID, state_archive_path: path, handoff_path: unit.handoffPath,
512
+ worker_session_id: unit.directExecution ? null : unit.childSessionID,
513
+ ...(unit.directExecution ? { execution_mode: "direct", executor_session_id: unit.directExecution.actor } : {}),
514
+ state_archive_path: path, handoff_path: unit.handoffPath,
512
515
  ...(anchor ? { accepted_anchor: anchor } : {}), evidence_id: proof.evidence_id,
513
516
  command: proof.execution.command, exit: proof.execution.exit_code, outcome: proof.execution.outcome,
514
517
  started_at: proof.execution.started_at, ended_at: proof.execution.ended_at,
@@ -534,7 +537,8 @@ export async function missionAcceptanceSummary(mission, run, operators, observe)
534
537
  try {
535
538
  if (!observe || !unit.childSessionID)
536
539
  throw new Error("native-worker-history-unavailable");
537
- return { ...provenance, status: "available", observations: await observe(unit.unit.validation, unit.childSessionID, unit.reviewerCorrection ? Date.parse(unit.reviewerCorrection.admittedAt ?? item.state.createdAt) : undefined) };
540
+ return { ...provenance, status: "available", observations: await observe(unit.unit.validation, unit.childSessionID, unit.directExecution ? Date.parse(unit.directExecution.startedAt) :
541
+ unit.reviewerCorrection ? Date.parse(unit.reviewerCorrection.admittedAt ?? item.state.createdAt) : undefined) };
538
542
  }
539
543
  catch (error) {
540
544
  return { ...provenance, status: "unavailable", reason: error instanceof Error ? error.message : String(error) };
@@ -554,8 +558,9 @@ export async function missionAcceptanceSummary(mission, run, operators, observe)
554
558
  verdict: mission.review.verdict, result: mission.review.result ?? null, self_recheck: mission.review.selfRecheck ?? null,
555
559
  freshness: "not established by run ID; existing source comparison remains required" } : null,
556
560
  delivery: { submission: mission.submission, git_lifecycle: current?.gitLifecycle ?? null,
561
+ observation: mission.deliveryObservation && mission.deliveryObservation.run_id === current?.runID ? mission.deliveryObservation : null,
557
562
  observation_source: "persisted operator Git lifecycle and formal validation records; no new Git inspection",
558
- clean: "not independently observed by this projection" },
563
+ clean: mission.deliveryObservation && mission.deliveryObservation.run_id === current?.runID ? mission.deliveryObservation.clean : "not independently observed by this projection" },
559
564
  interpretation: "Compare original requests with the submitted candidate and actual evidence. Historical PASS is not current PASS. Inspect concrete gaps, not routine archive searches or full source rereads. Existing completion and Review guards still apply." };
560
565
  }
561
566
  export function missionReviewIndependent(mission, child) {
@@ -575,7 +580,8 @@ export function missionPacket(mission, run) {
575
580
  note: "Historical results and spend are retained; they do not complete the current requirements." } } : {}),
576
581
  requirements: mission.requirements, original_request_refs: mission.requests.map(item => `user:${item.id}`),
577
582
  launch_conditions: mission.launchConditions ?? [], prohibited_write: mission.prohibitedWrite ?? [],
578
- accounting_scope: "Implementation units (Worker or scoped Reviewer correction) are not benchmark attempts. Host budget uses the existing Worker-unit ledger; orchestration, read-only Review and external campaign costs are excluded. Launch caps are fixed conditions, not a known campaign remainder.",
583
+ accounting_scope: "Implementation units (Worker, direct controller execution or scoped Reviewer correction) are not benchmark attempts. Host budget uses the existing unit ledger; orchestration, read-only Review and external campaign costs are excluded. Launch caps are fixed conditions, not a known campaign remainder.",
584
+ delivery_observation: mission.deliveryObservation?.run_id === (run?.runID ?? mission.runID) ? mission.deliveryObservation : null,
579
585
  submission: mission.submission, progress: mission.progress, consultations: mission.consultations ?? [],
580
586
  attempts: mission.attempts ?? [], corrections: (mission.corrections ?? []).map(({ findings: _findings, initialPrompt: _prompt, ...item }) => item), ...(mission.rescue ? { rescue: mission.rescue } : {}),
581
587
  operation: { kind: mission.kind ?? "implementation", status: missionExecutionStatus(mission),
@@ -123,6 +123,13 @@ interface UnitState {
123
123
  checks?: import("../plugin/runtime-bridge.js").ReviewerCorrectionCheck[];
124
124
  };
125
125
  evidence: readonly GoalEvidence[];
126
+ /** Implementation in the existing controller session, without a native child Task. */
127
+ directExecution?: {
128
+ actor: string;
129
+ startedAt: string;
130
+ finishedAt?: string;
131
+ checks: import("../plugin/runtime-bridge.js").ReviewerCorrectionCheck[];
132
+ };
126
133
  resultClass: string | null;
127
134
  failure?: SerialDispatchSettlement["failure"];
128
135
  normalRemediationUsed?: boolean;
@@ -345,6 +352,9 @@ export declare class OperatorRuntime {
345
352
  next(root: string, actor: string): Promise<unknown>;
346
353
  private nextOnce;
347
354
  admitWorker(root: string, actor: string, callID: string, args: unknown): Promise<OperatorTask>;
355
+ admitDirect(root: string, actor: string, callID: string): Promise<OperatorState>;
356
+ /** Continue an admitted independent Review in-place; no second native Task or prompt. */
357
+ admitReviewerDirect(root: string, author: string, callID: string, promptID: string): Promise<OperatorState>;
348
358
  private admitWorkerOnce;
349
359
  rejectedAdmission(root: string, callID: string, reason?: string): Promise<void>;
350
360
  rejectDispatch(root: string, callID: string, decision?: string): Promise<void>;