axstack 0.20.30 → 0.20.31

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,108 @@
1
+ # Autopilot
2
+
3
+ This reference applies to authorized engineering-delivery runs. The original
4
+ driver remains the sole run-record writer and phase router in the same chat;
5
+ phase completion is not a native ownership handoff. Explicit planning-only,
6
+ read-only, stop-after-phase, observation-only, and peer requests retain their
7
+ selected boundary. A status question such as "what's left" is observation,
8
+ not a mode change.
9
+
10
+ ## Advance and hold
11
+
12
+ Advance only after the finishing phase returns its completed identity (a
13
+ small-change intent, approved spec, matching ticket map, merge-ready or merged
14
+ state) and the run record has no open hold. A hold from any phase stops the run:
15
+ record its reason, owner, and resume condition, then take no dependent action.
16
+ That covers tracker access, adviser or arena-seat availability, diligence
17
+ FINDINGS when the phase records a hold, CI-wait timeout, readiness UNKNOWN,
18
+ dismissed approval, wake or cleanup uncertainty, single-provider routing, an
19
+ existing tag or version, and failed publish. Diligence FINDINGS during implement
20
+ follow its §6 repair route; at spec, tickets, or release preparation the driver
21
+ resolves them before advancing, and only a recorded hold pauses autopilot.
22
+
23
+ Record `Autopilot: on | paused (<hold>; resume: <condition>) | off (cancelled
24
+ <ts>)` and the next step in the private run record. A user answer to the hold
25
+ resumes after reconciliation; silence does not.
26
+ Awaiting human spec approval records `Autopilot: paused (spec approval; resume:
27
+ human approval)` as a decision hold eligible under the Notification policy.
28
+
29
+ ## Phase sequence
30
+
31
+ - Small: Align read-back, small-change intent, implement, watch in maintain
32
+ mode, human merge. An opted-in Align refinement is part of read-back.
33
+ - Substantial: Align, spec draft with advisers and diligence, human spec
34
+ approval at gate 1, tickets with diligence, implement, watch in maintain mode,
35
+ merge-ready, human merge. An opted-in Align refinement is part of gate 1.
36
+
37
+ Do not seek another phase-start instruction after a completed identity.
38
+ Spec approval is always the human's decision. Every PR merge is the human's,
39
+ including a release PR and each PR in a stack, bottom-up.
40
+
41
+ ## Implement into maintain watch
42
+
43
+ When implement publishes the run's first PR, arm exactly one `axstack-watch`
44
+ chat-run in authorized maintain mode. Use the 10-minute harness wake, with the
45
+ existing Orca fallback when unavailable. Later run PRs join after verified
46
+ publication readback; an explicitly adopted PR joins only with its maintenance
47
+ snapshot. The original driver alone routes work; one author writes each
48
+ candidate. Until a PR is merge-ready, wakes feed implement §6 step 4. After
49
+ merge-ready, watch §5 maintenance repairs feedback, rebases when the base moves,
50
+ keeps CI green, and checks approvals without re-requesting human review.
51
+
52
+ Maintain is the default mode for run-created PRs. End the chat-run watch when
53
+ every watched PR is merged or closed and the run's release step is settled or
54
+ not applicable, or when the user cancels. Expiry is a recorded stop with
55
+ resumable state, never a silent renewal. A required PR closed without merging
56
+ is incomplete scope; it does not make the run release-eligible. On wake expiry
57
+ record `Autopilot: paused (wake expired; resume: user reauthorizes a wake)` and
58
+ notify under the recorded Notification policy when user action is needed.
59
+
60
+ ## Release and install, when applicable
61
+
62
+ Detect applicability once at Align or spec time. Record `Release: <AGENTS.md
63
+ file:line + tag-triggered workflow path + named install hosts> | not applicable
64
+ (<reason>)`. The predicate is an AGENTS.md release rule naming an existing
65
+ tag-triggered workflow. A partial match is not applicable and its reason is
66
+ noted. Install hosts come only from explicit targets; an absent host list is a
67
+ decision hold, not permission to infer hosts. A missing install host list at
68
+ Align or spec time is a decision hold before release authority is presented.
69
+
70
+ Show the `Release:` line in the spec for human approval at gate 1, or the small
71
+ work Align read-back. Copy that decision to `Authority:` in the run record.
72
+ This authority is per run and never carries over to another run or repository.
73
+ The small-work Align read-back names the existing Release and host-mutation
74
+ authority and explicit hosts; silence cannot fill a missing authority or target.
75
+
76
+ After all required feature PRs merge, open one release PR. Default to a patch
77
+ version, or minor if a `feat` commit landed since the last tag. This normal run
78
+ PR gets authored review and diligence of its body against merged PRs, reaches
79
+ merge-ready, then waits for human merge. Once the forge confirms that merge,
80
+ tag and wait for the staged publish. Human npm stage approval is a decision
81
+ hold: agents never run `npm stage approve`. A wake verifies the registry reports
82
+ the expected package and version. Install on the named hosts, verify version
83
+ and roles, then run Close-out last with release and install receipts and the
84
+ installed version.
85
+
86
+ An existing version or tag, failed publish, pending approval, uncertain
87
+ registry result, missing host access, or failed install verification is a
88
+ resumable hold, never success. Tagging, publishing, installation, and host
89
+ mutation require the recorded per-run authority and their existing checks.
90
+
91
+ ## Resume, cancel, and notify
92
+
93
+ At every entry (user message, wake, compaction, or new chat), reconcile the
94
+ owner, authoritative Dispatch, approved revision, PR membership, uncertain
95
+ tags, wakes, publications, and completed receipts under lifecycle and
96
+ run-record before advancing. Only the original driver advances. Wakes do not
97
+ reset attempt budgets and do not grant approvals. Cancel sets `Autopilot: off`,
98
+ stops new actions, and ends the watch under watch §6 with guarded settlement.
99
+ Cancellation does not cancel a running author Dispatch by inference; let it
100
+ report, then settle that exact Dispatch under lifecycle guards without new
101
+ publication.
102
+
103
+ Use the run's recorded Notification policy through `axstack-relay`.
104
+ Decision holds, including spec and npm approval, are always eligible. Across
105
+ implementation and release, merge-ready and merged notifications together are
106
+ capped at two per run; deduplicate by purpose and revision. Healthy ticks stay
107
+ quiet. A failed or uncertain delivery preserves the underlying hold. A relay
108
+ message is only a notification, never authority to approve, merge, or publish.
@@ -40,7 +40,11 @@ Validate the configured provider and model at actual launch. If it is
40
40
  unavailable or exhausted, pause affected work, record the gap, and ask the
41
41
  user. Never infer a route from quota state or subscription entitlement. Every
42
42
  substitution requires the user's decision: configured alternatives and native
43
- fallback prose are not defaults.
43
+ fallback prose are not defaults. The only within-class exception is explicit
44
+ model rejection before the first turn: Codex may retry with `--retry-of` using
45
+ the next eligible version in the same class, provider, and effort, recording
46
+ the failed ID, error, and fallback ID. Claude rejection holds. Timeout, quota,
47
+ and auth failures hold.
44
48
 
45
49
  ## Driver and adviser split
46
50
 
@@ -38,23 +38,42 @@ Read `roles.json` from the installed shared root `skills/axstack/`. The installe
38
38
  shape is `{ "version": 1, "preset": "<name>", "roles": [...] }`. Bundled
39
39
  profiles are setup inputs shaped as
40
40
  `{ "version": 1, "roles": [...] }`. A new run records the selected preset and
41
- all 32 role rows once. An active run keeps the exact snapshot until the user
42
- explicitly changes it.
43
-
44
- Select the requested role by stable ID. A missing or null model holds only that role;
45
- never launch a provider default. Launch-by-agent-id routes for which Orca exposes no
41
+ all 32 role rows once. For each role record class, resolved exact ID, source,
42
+ and time. An active run keeps the exact snapshot; resume reuses it without
43
+ re-resolution until the user explicitly changes it.
44
+
45
+ Select the requested role by stable ID. A missing class and missing or null
46
+ model holds only that role; never launch a provider default. Resolve Codex
47
+ classes with `scripts/resolve-models.js`, passing the catalog path explicitly;
48
+ missing or malformed catalogs hold. The first launch of each Claude class uses
49
+ its alias. Read the exact ID from the first assistant turn's `message.model` in
50
+ that worker's own session transcript at
51
+ `~/.claude/projects/<worktree-path-slug>/*.jsonl`; the worktree path slug
52
+ replaces each non-alphanumeric character with `-`. Identify the file by the
53
+ worker's session ID, or use the newest file created after launch. Later launches
54
+ of that class use the recorded exact ID. Before read-back record `alias,
55
+ unresolved`; record an unknown read-back as unknown and hold
56
+ provenance-dependent work. A worker self-report is a labeled last
57
+ resort. Launch-by-agent-id routes for which Orca exposes no
46
58
  `--model` override (today: `grok`, `antigravity`) record `model: null` with an explicit note and are
47
59
  launchable; the run record snapshots the model the TUI reports. Validate provider, model, and effort
48
60
  against the guide and actual launch capability. Stored `modeId` and other
49
61
  permission fields are conservative intent, not proof of effective permission
50
62
  parity or a security boundary. Requested settings, input acceptance, effective
51
63
  settings, and completed work are separate evidence. An unsupported or
52
- unavailable value holds affected work for the user's decision without fallback.
64
+ unavailable value holds affected work for the user's decision except the narrow
65
+ retry below.
53
66
  The single-provider preset's null adviser and round-2 seat are intentional installation data, not
54
67
  readiness failure; because Align and Spec require both adviser receipts, either
55
68
  null adviser still holds those phases. The current chat is the driver and has
56
69
  no role row in any preset.
57
70
 
71
+ Only explicit model rejection before the first turn permits a Codex
72
+ `--retry-of` with the next eligible ID in the same class, provider, and effort.
73
+ Fence the rejected Dispatch and record tried ID, error, and fallback ID in the
74
+ snapshot and reply. Timeout, quota, auth, and other failures hold; Claude
75
+ rejection holds. Apply this to every role, including advisers and judges.
76
+
58
77
  ## Materialize checkouts as worktrees of the registered repo
59
78
 
60
79
  Every reviewer, release, or worker checkout is `ORCA worktree create --repo
@@ -5,12 +5,12 @@
5
5
  - `axstack-reviewer-primary` and `axstack-reviewer-secondary` are the ordered
6
6
  peer pair. Peer review uses both; authored review uses this table:
7
7
 
8
- | Preset | Author | Reviewer (model/effort) |
8
+ | Preset | Author class | Reviewer (class/effort) |
9
9
  | --- | --- | --- |
10
- | `mixed` | Codex / Sol (`codex/gpt-6-sol`) | `axstack-reviewer-secondary` (`claude/claude-opus-5-5` medium) |
11
- | `mixed` | Claude / Opus (`claude/claude-opus-5-5`) | `axstack-reviewer-primary` (`codex/gpt-6-sol` high) |
12
- | `codex-only` | Codex / Sol (`codex/gpt-6-sol`) | `axstack-reviewer-secondary` (`codex/gpt-6-luna` xhigh) |
13
- | `claude-only` | Claude / Opus (`claude/claude-opus-5-5`) | `axstack-reviewer-secondary` (`claude/claude-sonnet-5-5` high) |
10
+ | `mixed` | `codex/sol` | `axstack-reviewer-secondary` (`claude/opus` medium) |
11
+ | `mixed` | `claude/opus` | `axstack-reviewer-primary` (`codex/sol` high) |
12
+ | `codex-only` | `codex/sol` | `axstack-reviewer-secondary` (`codex/luna` xhigh) |
13
+ | `claude-only` | `claude/opus` | `axstack-reviewer-secondary` (`claude/sonnet` high) |
14
14
  - `axstack-advisor-astra`/`axstack-advisor-opus` advise and author candidates;
15
15
  `axstack-arena-candidate-grok`/
16
16
  `axstack-arena-candidate-antigravity` add families.
@@ -24,11 +24,11 @@
24
24
  - `axstack-auditor`/`axstack-research-requirements`/
25
25
  `axstack-research-code`/`axstack-research-web`/
26
26
  `axstack-explore-execution`/`axstack-monitor`:
27
- `claude-sonnet-5-5` high in mixed/claude-only.
27
+ `claude/sonnet` high in mixed/claude-only.
28
28
  `axstack-monitor`: standalone watch never sends; chat-run watch: bounded
29
29
  internal reports to its Run and original driver.
30
30
  - Sol pairs `axstack-auditor-sol`/`axstack-research-code-sol`/
31
- `axstack-explore-execution-sol`: `codex/gpt-6-sol` high in
31
+ `axstack-explore-execution-sol`: `codex/sol` high in
32
32
  mixed/codex-only; intentionally absent in claude-only. Dispatch each
33
33
  independently from its Sonnet seat on the same bounded brief without
34
34
  cross-reading. The driver reconciles findings per claim, never averages.
@@ -11,22 +11,33 @@ skills root, or an explicit user selection in the run record. Missing or contrad
11
11
  a setup gap: hold. Never infer from live profiles or `list_profiles`, harness,
12
12
  tools, credentials, quota, subscription, or default to `mixed`.
13
13
 
14
- At start, snapshot all 32 role IDs with provider/model/mode/effort; absent
14
+ At start, snapshot all 32 role IDs with provider/modelClass/model/mode/effort; absent
15
15
  or unconfigured roles are recorded explicitly; never default.
16
16
  Such a role holds only its work. Later installed or changed roles need an
17
17
  explicit user decision to enter the snapshot. Live profiles
18
18
  are authoritative at snapshot time and for availability; bundled presets are setup
19
19
  inputs, not runtime proof.
20
+ For each role record class, resolved exact ID, source (catalog, transcript, or
21
+ pin), and time. Codex classes resolve through
22
+ `skills/axstack/scripts/resolve-models.js` with an explicit catalog
23
+ path; missing or malformed catalog holds. Claude classes start as `alias,
24
+ unresolved` until transcript read-back. Resume must reuse the snapshot and
25
+ never re-resolve it.
20
26
 
21
27
  Preset changes apply to new runs only; an active run keeps its snapshot.
22
28
  Changing it or replacing a session needs an explicit user decision and
23
29
  revalidation. Unavailable models, efforts, roles, or overrides hold only affected
24
30
  work; no automatic fallback, quota routing, subscription inference, or silent
25
- provider/model/effort substitution.
31
+ provider/model/effort substitution. Only
32
+ explicit model rejection before the first turn permits Codex `--retry-of` with
33
+ the next eligible ID in the same class, provider, and effort. Fence the failed
34
+ Dispatch and record tried ID, error, and fallback ID in the snapshot and reply.
35
+ Timeout, quota, auth, and other failures hold; Claude rejection holds.
26
36
 
27
37
  Load the [Role roster](role-roster.md) for configured roles and authored-review pairings.
28
38
 
29
- Provenance is matched on provider/model ID; effort never maps. Missing table-row
39
+ Provenance is matched on provider/model class derived from the recorded exact ID;
40
+ effort never maps. Missing table-row
30
41
  provenance is unsupported and `INCOMPLETE`; report it and ask the user. Never
31
42
  infer from slot, driver, owner, or provider. Author and owner never review.
32
43
 
@@ -75,7 +86,7 @@ reason in the run record, or in the brief for tiny direct work.
75
86
  Require an approved spec plus a ticket map tied to that exact spec
76
87
  revision, with acceptance checks and dependencies in the explicitly selected
77
88
  Markdown, GitHub Issues, or Linear store. Prepare via `axstack-align` -> `axstack-spec`
78
- (one approval) -> `axstack-tickets` -> handoff, then stop.
89
+ (one approval) -> `axstack-tickets` -> handoff, then continue under autopilot when eligible.
79
90
  - **Small:** clear, bounded one-PR work. The driver captures the named
80
91
  **small-change intent** from the current request or user-chosen existing
81
92
  issue plus explicit acceptance checks and exclusions, snapshots it once, and
@@ -96,7 +107,7 @@ not alone a formal spec trigger. Hold affected unsafe work while reassessing.
96
107
  ## Lifecycle routes (mode-specific scope identity required)
97
108
 
98
109
  - Preparation: substantial work follows the align -> spec -> tickets ->
99
- handoff path above, then stops; small work uses the driver-captured
110
+ handoff path above, then continues under autopilot when eligible; small work uses the driver-captured
100
111
  small-change intent.
101
112
  - Execution: with its identity present, `axstack-implement` ->
102
113
  `axstack-review` -> `axstack-watch`.
@@ -151,6 +151,8 @@ Authority: <who authorized which mutation>
151
151
  Intent: <approved spec rev | small-change intent | adopted snapshot | peer/read-only mode>
152
152
  Routing: <preset + source + snapshot ref>
153
153
  Notification policy: <none | transport/target label/host/instructions path>
154
+ Autopilot: on | paused (<hold>; resume: <condition>) | off (cancelled <ts>); next: <step>
155
+ Release: <AGENTS.md file:line + tag-triggered workflow path + named install hosts> | not applicable (<reason>)
154
156
  Source base: <exact revision or source identity>
155
157
  IDs: <repo/project + workspace/agent receipt pointers>
156
158
  Worktrees in other repositories: <per-run repository and worktree IDs or none>
@@ -0,0 +1,46 @@
1
+ import { readFileSync } from 'node:fs';
2
+
3
+ const args = process.argv.slice(2);
4
+ const option = (name) => {
5
+ const index = args.indexOf(name);
6
+ return index < 0 ? null : args[index + 1];
7
+ };
8
+ const path = option('--catalog');
9
+ const modelClass = option('--class');
10
+ const effort = option('--effort');
11
+
12
+ try {
13
+ if (!path || !['astra', 'sol', 'luna'].includes(modelClass) || !effort) {
14
+ throw new Error('expected --catalog path --class astra|sol|luna --effort level');
15
+ }
16
+ const catalog = JSON.parse(readFileSync(path, 'utf8'));
17
+ if (!Array.isArray(catalog.models) || typeof catalog.client_version !== 'string'
18
+ || typeof catalog.fetched_at !== 'string') {
19
+ throw new Error('malformed catalog');
20
+ }
21
+ const classPattern = new RegExp(`^gpt-(\\d+(?:\\.\\d+)*)-${modelClass}$`);
22
+ const candidates = catalog.models
23
+ .filter((entry) => entry && entry.visibility === 'list'
24
+ && typeof entry.slug === 'string'
25
+ && classPattern.test(entry.slug)
26
+ && Array.isArray(entry.supported_reasoning_levels)
27
+ && entry.supported_reasoning_levels.some((level) => level?.effort === effort))
28
+ .map((entry) => entry.slug)
29
+ .sort((left, right) => {
30
+ const a = left.slice(4, -(modelClass.length + 1)).split('.').map(Number);
31
+ const b = right.slice(4, -(modelClass.length + 1)).split('.').map(Number);
32
+ for (let i = 0; i < Math.max(a.length, b.length); i++) {
33
+ const difference = (b[i] ?? 0) - (a[i] ?? 0);
34
+ if (difference) return difference;
35
+ }
36
+ return 0;
37
+ });
38
+ if (!candidates.length) throw new Error(`no eligible ${modelClass} model in catalog`);
39
+ console.log(JSON.stringify({
40
+ provider: 'codex', modelClass, effort, model: candidates[0], candidates,
41
+ catalog: { path, client_version: catalog.client_version, fetched_at: catalog.fetched_at },
42
+ }));
43
+ } catch (error) {
44
+ console.error(`model catalog resolution hold: ${error.message}`);
45
+ process.exitCode = 1;
46
+ }
@@ -5,6 +5,9 @@ description: When exploring or planning engineering work, use axstack-align to s
5
5
 
6
6
  # Align
7
7
 
8
+ For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
9
+ for phase continuation and holds.
10
+
8
11
  On driver entry, sweep under [Workspace hygiene](../axstack/references/workspace-hygiene.md); dispatched workers do not sweep.
9
12
  For every dispatch brief, name its private `<run dir>/evidence/<dispatch>/` folder.
10
13
 
@@ -166,7 +169,7 @@ record or spec. Read-only scope keeps proposed documentation in the permitted
166
169
  private record or response. Documentation is neither implementation nor spec
167
170
  approval; record chosen document names and paths once per run.
168
171
 
169
- ## Read back, classify, and stop
172
+ ## Read back, classify, and route
170
173
 
171
174
  1. Read back the decisions, constraints, exclusions, and remaining evidence
172
175
  gaps. For substantial work, this summary becomes part of the draft spec in
@@ -189,6 +192,7 @@ approval; record chosen document names and paths once per run.
189
192
  [Orca runtime](../axstack/references/orca-runtime.md) immediately before
190
193
  actual dispatch. Alignment completion never dispatches a recipient.
191
194
 
192
- Alignment stops for both sizes only when the handoff is usable, its next scope
193
- identity is explicit, and execution has not started. The user invokes
194
- `axstack-implement` to execute.
195
+ Alignment completes for both sizes only when the handoff is usable and its next
196
+ scope identity is explicit. An eligible delivery run continues under Autopilot;
197
+ an explicit stop-after-Align request ends here. Substantial work continues to
198
+ Spec, and small work continues from its small-change intent to Implement.
@@ -28,8 +28,8 @@ immediately before an actual auditor profile or session dispatch. Ordinary
28
28
  audit reading and record writing do not load it, and the auditor never
29
29
  dispatches.
30
30
 
31
- Core owns the `axstack-auditor` profile (claude/claude-sonnet-5-5 high in
32
- mixed/claude-only; codex/gpt-6-luna xhigh in codex-only) and its invocation.
31
+ Core owns the `axstack-auditor` profile (claude/sonnet high in
32
+ mixed/claude-only; codex/luna xhigh in codex-only) and its invocation.
33
33
  This skill governs what that auditor reads, measures, and proposes.
34
34
  Dispatch `axstack-auditor` and `axstack-auditor-sol` independently on the same
35
35
  bounded brief, without cross-reading. The driver reconciles findings per claim;
@@ -91,6 +91,14 @@ counts with denominators plus the evidence behind the count:
91
91
  evidence path is absent, or `UNKNOWN` with the reason when its records are
92
92
  unavailable.
93
93
  - Independent exact-revision review status and unresolved findings.
94
+ For authored PRs, measure the selected reviewer against this class pairing:
95
+
96
+ | Preset | Author class | Reviewer (class/effort) |
97
+ | --- | --- | --- |
98
+ | `mixed` | `codex/sol` | `axstack-reviewer-secondary` (`claude/opus` medium) |
99
+ | `mixed` | `claude/opus` | `axstack-reviewer-primary` (`codex/sol` high) |
100
+ | `codex-only` | `codex/sol` | `axstack-reviewer-secondary` (`codex/luna` xhigh) |
101
+ | `claude-only` | `claude/opus` | `axstack-reviewer-secondary` (`claude/sonnet` high) |
94
102
  - Simplification applicability determinations evidenced / total candidate diffs,
95
103
  and complete simplification receipts / total candidates, broken down as
96
104
  `applied`, `not-applicable`, or `UNKNOWN` with the reason. This measures
@@ -5,6 +5,9 @@ description: When an approved task is ready to build or repair, use axstack-impl
5
5
 
6
6
  # Implement
7
7
 
8
+ For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
9
+ for phase continuation and holds.
10
+
8
11
  On driver entry, sweep under [Workspace hygiene](../axstack/references/workspace-hygiene.md); dispatched workers do not sweep.
9
12
  For every dispatch brief, name its private `<run dir>/evidence/<dispatch>/` folder.
10
13
  Include [Safe deletion](../axstack/references/workspace-hygiene.md#safe-deletion) in author briefs.
@@ -215,23 +218,34 @@ For each PR:
215
218
  `repairs` once and counts once toward the third-round hold.
216
219
 
217
220
  One run-level completion wait covers every unsettled Dispatch; the bounded
218
- forge check wait is the only other wait. End a turn only when every required PR
221
+ forge check wait is the only other implementation wait. The eligible run arms
222
+ one maintain-mode chat-run watch at its first published PR; that watch owns its
223
+ 10-minute harness wake or Orca fallback. End a turn only when every required PR
219
224
  is `merge-ready` or `held`. Under the recorded Notification policy,
220
- `axstack-relay` sends only a serious risk immediately or a genuine blocked
221
- operation that needs user intervention after bounded safe recovery. Questions,
222
- spec approvals, progress, CI pending, merge-ready, merged, and completion stay
225
+ `axstack-relay` sends only a serious risk immediately, a genuine blocked
226
+ operation needing user intervention after bounded safe recovery, or the
227
+ decision holds and capped milestones named by the recorded Notification policy.
228
+ Routine questions stay in Orca. Progress, CI pending, and completion always stay
223
229
  in Orca.
230
+ Only the bounded categories—user-decision holds (including spec approval),
231
+ serious-risk holds, and at most two merge-ready/merged milestones per run—may
232
+ be relayed under the recorded Notification policy.
224
233
 
225
234
  Merge-ready is the human boundary: the user merges, bottom-up for a stack. The
226
- driver resumes on the user's next message or `/axstack-watch`; no Orca merge
227
- wake exists today. Re-read forge state: record forge-merged PRs as `merged`;
235
+ driver resumes on the user's next message, `/axstack-watch`, or the armed
236
+ chat-run watch wake; no Orca merge wake exists today.
237
+ Re-read forge state: record forge-merged PRs as `merged`;
228
238
  changed heads or feedback return to step 1; retain useful author work before Close-out.
229
- Run Close-out once only after every required PR is forge-merged and acceptance
230
- passes. It settles workers, records counts, makes the auditor decision and
239
+ Run Close-out once only after every required PR is forge-merged, the run's
240
+ Release step is settled or not applicable, and acceptance passes. It settles
241
+ workers, records counts, makes the auditor decision and
231
242
  settlement, releases worktrees, closes eligible tickets, and archives the run.
243
+ Without an Autopilot or Release record, the Release step is not applicable for
244
+ both Close-out and run completion.
232
245
 
233
246
  The loop requires the `mixed` two-provider authored-review row. `codex-only` or
234
247
  `claude-only` holds at step (3) for an explicit user routing choice, with no
235
248
  substitution or same-provider review. Derived PR states are `authoring |
236
249
  published | in-review | repairing(n) | merge-ready | merged | held`. The run is
237
- done only when every required PR is forge-merged and Close-out has receipts.
250
+ done only when every required PR is forge-merged, the run's Release step is
251
+ settled or not applicable, and Close-out has receipts.
@@ -5,6 +5,9 @@ description: When the user requests a relay message or test, or an authorized no
5
5
 
6
6
  # Relay
7
7
 
8
+ For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
9
+ for phase continuation and holds.
10
+
8
11
  Send normal messages, transport tests, and authorized notifications to the
9
12
  user through Hermes' native one-way `hermes send`. This is an inline caller
10
13
  procedure: it creates no driver, team, owner, auditor, monitor, child session,
@@ -25,8 +28,12 @@ Choose the applicable message type:
25
28
  in the caller's private notification policy. State the issue, impact, and the
26
29
  answer or action needed.
27
30
  - **Routine run events:** questions, spec approvals, progress, CI pending,
28
- merge-ready, merged, and completion stay in Orca. They never become proactive
29
- relay messages merely because the run is waiting.
31
+ merge-ready, merged, and completion stay in Orca unless the recorded
32
+ Notification policy names it. A policy may name only user-decision holds and
33
+ at most two merge-ready/merged milestones per run; deduplicate across implementation
34
+ and release. Progress, CI pending, and completion are never eligible merely
35
+ because a policy exists. They never become proactive relay messages merely
36
+ because the run is waiting.
30
37
 
31
38
  Verify the transport, execution host, and intended recipient from the user's
32
39
  request, trusted caller context, or an existing private notification policy.
@@ -204,17 +204,18 @@ This section applies to peer and authored PR modes.
204
204
  - **Authored:** exactly one eligible independent reviewer from this complete
205
205
  mapping:
206
206
 
207
- | Preset | Actual author provider/model | Reviewer role (configured model/effort) |
207
+ | Preset | Actual author provider/class | Reviewer role (configured class/effort) |
208
208
  | --- | --- | --- |
209
- | `mixed` | Codex / Sol (`codex/gpt-6-sol`) | `axstack-reviewer-secondary` (`claude/claude-opus-5-5` medium) |
210
- | `mixed` | Claude / Opus (`claude/claude-opus-5-5`) | `axstack-reviewer-primary` (`codex/gpt-6-sol` high) |
211
- | `codex-only` | Codex / Sol (`codex/gpt-6-sol`) | `axstack-reviewer-secondary` (`codex/gpt-6-luna` xhigh) |
212
- | `claude-only` | Claude / Opus (`claude/claude-opus-5-5`) | `axstack-reviewer-secondary` (`claude/claude-sonnet-5-5` high) |
209
+ | `mixed` | `codex/sol` | `axstack-reviewer-secondary` (`claude/opus` medium) |
210
+ | `mixed` | `claude/opus` | `axstack-reviewer-primary` (`codex/sol` high) |
211
+ | `codex-only` | `codex/sol` | `axstack-reviewer-secondary` (`codex/luna` xhigh) |
212
+ | `claude-only` | `claude/opus` | `axstack-reviewer-secondary` (`claude/sonnet` high) |
213
213
 
214
214
  The diligence receipt is separate and does not count as a reviewer receipt.
215
215
 
216
- Provenance is matched on provider/model ID; record effort, but never use
217
- effort to create a mapping. Any other author provenance for the
216
+ From the recorded exact model ID, derive its class and match provenance
217
+ on provider/class; record effort, but never use effort to create a mapping.
218
+ An ID with no class is `INCOMPLETE`. Any other author provenance for the
218
219
  selected preset is unsupported and `INCOMPLETE`, including its secondary
219
220
  reviewer model, Astra, Luna, or Fable. Report the exact provenance gap and
220
221
  ask the user. Never derive a reverse pairing from slot position. The
@@ -233,6 +234,11 @@ This section applies to peer and authored PR modes.
233
234
  effort and spawn no redundant final reviewer. If a required reviewer is
234
235
  unavailable, report that exact model gap, mark review `INCOMPLETE`, and ask
235
236
  the user; do not lower effort or choose any automatic fallback.
237
+ The only within-class exception is explicit model rejection before the first
238
+ turn: Codex may use
239
+ `--retry-of` with the next eligible ID in the same class, provider, and
240
+ effort; fence and record the rejected attempt. Timeout, quota, and auth
241
+ failures hold; Claude rejection holds. Never cross class or provider.
236
242
 
237
243
  Continue only when session receipts prove the required models, non-author
238
244
  independence, actual author provenance where applicable, and exact brief.
@@ -5,6 +5,9 @@ description: When agreed work needs an approved baseline, use axstack-spec to wr
5
5
 
6
6
  # Specification baseline
7
7
 
8
+ For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
9
+ for phase continuation and holds.
10
+
8
11
  On driver entry, sweep under [Workspace hygiene](../axstack/references/workspace-hygiene.md); dispatched workers do not sweep.
9
12
  For every dispatch brief, name its private `<run dir>/evidence/<dispatch>/` folder.
10
13
 
@@ -75,7 +78,8 @@ Material change: <none | description + affected PRs/tasks + hold state>
75
78
  The snapshot is ready for ticketing when its authoritative revision,
76
79
  counterpart, and preserved ref resolve to the approved content. Return that
77
80
  exact identity; routine execution of the settled plan needs no repeat adviser
78
- consultation or spec approval.
81
+ consultation or spec approval. In an eligible delivery run with no hold,
82
+ continue to Tickets in the same driver chat.
79
83
 
80
84
  ## Material revisions
81
85
 
@@ -5,12 +5,15 @@ description: When an approved capability needs executable tasks, use axstack-tic
5
5
 
6
6
  # Tickets
7
7
 
8
+ For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
9
+ for phase continuation and holds.
10
+
8
11
  On driver entry, sweep under [Workspace hygiene](../axstack/references/workspace-hygiene.md); dispatched workers do not sweep.
9
12
  For every dispatch brief, name its private `<run dir>/evidence/<dispatch>/` folder.
10
13
 
11
14
  Produce an executable capability map tied to the exact approved spec revision.
12
15
  Keep user-visible capabilities in the selected store, keep implementation detail
13
- in the repository, reconcile lifecycle state, and stop before implementation.
16
+ in the repository, reconcile lifecycle state, and return a map for continuation.
14
17
 
15
18
  Before mapping, load [Standing contracts](../axstack/references/contracts.md).
16
19
  Follow its required edge to [Shared lifecycle](../axstack/references/lifecycle.md),
@@ -100,5 +103,5 @@ Recommendation: <move to In Review | keep open | close | other> (driver verifies
100
103
 
101
104
  5. **Return the mapping.** Report the pinned spec revision, selected store, map
102
105
  references, mutations performed by the driver, recorded gaps, and unresolved
103
- decisions. Stop with a map ready for lifecycle continuation; implementation
104
- has not started.
106
+ decisions. With a complete map and no hold, an eligible delivery run
107
+ continues to Implement in the same driver chat.
@@ -5,6 +5,9 @@ description: When babysitting an existing PR, use axstack-watch to monitor or ma
5
5
 
6
6
  # Watch
7
7
 
8
+ For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
9
+ for phase continuation and holds.
10
+
8
11
  On driver entry, sweep under [Workspace hygiene](../axstack/references/workspace-hygiene.md); dispatched workers do not sweep.
9
12
  For every dispatch brief, name its private `<run dir>/evidence/<dispatch>/` folder.
10
13
 
@@ -49,7 +52,8 @@ authority is unverified, record the hold and continue read-only.
49
52
 
50
53
  Choose one mode from the user's authority and record it before dispatch:
51
54
 
52
- - **Chat-run watch:** the initiating chat remains the only driver and record
55
+ - **Chat-run watch:** authorized maintain mode is the default for run-created
56
+ PRs. The initiating chat remains the only driver and record
53
57
  writer for every PR raised in its Run, including later verified publications
54
58
  and explicitly adopted members. Follow [Chat-run watch runtime](references/watch-runtime.md#chat-run-watch)
55
59
  for its scheduled driver wake and Orca fallback. This mode has no replacement `axstack-owner` or
@@ -121,11 +125,16 @@ A handled wake has an acknowledged event ID, an observation or action bound to
121
125
  the current revision, and a recorded hold or next owner where work remains.
122
126
 
123
127
  Under a recorded `Notification policy`, the owner may use the optional
124
- [axstack-relay](../axstack-relay/SKILL.md) only for a serious risk immediately
125
- or a genuine blocked operation needing user intervention after bounded safe
126
- recovery. Questions, spec approvals, progress, CI pending, merge-ready, merged,
127
- and completion stay in Orca. The standalone monitor never sends; the chat-run
128
- observer reports only internally. Deduplicate authorized notifications;
128
+ [axstack-relay](../axstack-relay/SKILL.md) only for a serious risk immediately,
129
+ a genuine blocked operation needing user intervention after bounded safe
130
+ recovery, or decision holds and capped milestones named by the recorded policy.
131
+ Routine questions stay in Orca. Progress, CI pending, and completion always stay
132
+ in Orca.
133
+ Only the bounded categories—user-decision holds (including spec approval),
134
+ serious-risk holds, and at most two merge-ready/merged milestones per run—may
135
+ be relayed under the recorded Notification policy.
136
+ The standalone monitor never sends; the chat-run observer reports only
137
+ internally. Deduplicate authorized notifications;
129
138
  absent policy or failed relay uses the current Orca conversation and leaves
130
139
  the existing hold open.
131
140
 
@@ -148,10 +157,14 @@ human approval remain allowed.
148
157
 
149
158
  ## 6. End and preserve continuity
150
159
 
151
- End a chat-run watch after all members merged or closed, user cancellation, or
152
- the recorded wake expires. Stop the chosen wake and verify its stop receipt;
153
- a failed or uncertain harness wake stop is a hold.
154
- the Orca fallback also needs own-automation disable/readback and driver-owned automation
160
+ End a chat-run watch after all members merged or closed and the run's release
161
+ step is settled or not applicable, user cancellation, or the recorded wake
162
+ expires. Without an Autopilot or Release record, the release step is not
163
+ applicable to this watch. A required PR closed without merging records a
164
+ decision hold and the wake remains active while unexpired until the user
165
+ resolves scope, cancels, or the wake expires. Stop the chosen wake and verify
166
+ its stop receipt; a failed or uncertain harness wake stop is a hold.
167
+ The Orca fallback also needs own-automation disable/readback and driver-owned automation
155
168
  removal and workspace cleanup under
156
169
  [Watch runtime](references/watch-runtime.md#chat-run-watch).
157
170
 
@@ -183,6 +196,6 @@ Resume: <known commands or verified refs needed to reconcile from this revision>
183
196
 
184
197
  The watch ends only when registrations are stopped, receipts are recorded, and
185
198
  the PR is either merged or represented by this resumable state.
186
- When every required PR is merged, follow the lifecycle
187
- [Close-out](../axstack/references/lifecycle.md#close-out) before reporting the
188
- run as done.
199
+ When every required PR is merged and the run's Release step is settled or not
200
+ applicable, follow the lifecycle [Close-out](../axstack/references/lifecycle.md#close-out)
201
+ before reporting the run as done.