axstack 0.20.29 → 0.20.31

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,108 @@
1
+ # Autopilot
2
+
3
+ This reference applies to authorized engineering-delivery runs. The original
4
+ driver remains the sole run-record writer and phase router in the same chat;
5
+ phase completion is not a native ownership handoff. Explicit planning-only,
6
+ read-only, stop-after-phase, observation-only, and peer requests retain their
7
+ selected boundary. A status question such as "what's left" is observation,
8
+ not a mode change.
9
+
10
+ ## Advance and hold
11
+
12
+ Advance only after the finishing phase returns its completed identity (a
13
+ small-change intent, approved spec, matching ticket map, merge-ready or merged
14
+ state) and the run record has no open hold. A hold from any phase stops the run:
15
+ record its reason, owner, and resume condition, then take no dependent action.
16
+ That covers tracker access, adviser or arena-seat availability, diligence
17
+ FINDINGS when the phase records a hold, CI-wait timeout, readiness UNKNOWN,
18
+ dismissed approval, wake or cleanup uncertainty, single-provider routing, an
19
+ existing tag or version, and failed publish. Diligence FINDINGS during implement
20
+ follow its §6 repair route; at spec, tickets, or release preparation the driver
21
+ resolves them before advancing, and only a recorded hold pauses autopilot.
22
+
23
+ Record `Autopilot: on | paused (<hold>; resume: <condition>) | off (cancelled
24
+ <ts>)` and the next step in the private run record. A user answer to the hold
25
+ resumes after reconciliation; silence does not.
26
+ Awaiting human spec approval records `Autopilot: paused (spec approval; resume:
27
+ human approval)` as a decision hold eligible under the Notification policy.
28
+
29
+ ## Phase sequence
30
+
31
+ - Small: Align read-back, small-change intent, implement, watch in maintain
32
+ mode, human merge. An opted-in Align refinement is part of read-back.
33
+ - Substantial: Align, spec draft with advisers and diligence, human spec
34
+ approval at gate 1, tickets with diligence, implement, watch in maintain mode,
35
+ merge-ready, human merge. An opted-in Align refinement is part of gate 1.
36
+
37
+ Do not seek another phase-start instruction after a completed identity.
38
+ Spec approval is always the human's decision. Every PR merge is the human's,
39
+ including a release PR and each PR in a stack, bottom-up.
40
+
41
+ ## Implement into maintain watch
42
+
43
+ When implement publishes the run's first PR, arm exactly one `axstack-watch`
44
+ chat-run in authorized maintain mode. Use the 10-minute harness wake, with the
45
+ existing Orca fallback when unavailable. Later run PRs join after verified
46
+ publication readback; an explicitly adopted PR joins only with its maintenance
47
+ snapshot. The original driver alone routes work; one author writes each
48
+ candidate. Until a PR is merge-ready, wakes feed implement §6 step 4. After
49
+ merge-ready, watch §5 maintenance repairs feedback, rebases when the base moves,
50
+ keeps CI green, and checks approvals without re-requesting human review.
51
+
52
+ Maintain is the default mode for run-created PRs. End the chat-run watch when
53
+ every watched PR is merged or closed and the run's release step is settled or
54
+ not applicable, or when the user cancels. Expiry is a recorded stop with
55
+ resumable state, never a silent renewal. A required PR closed without merging
56
+ is incomplete scope; it does not make the run release-eligible. On wake expiry
57
+ record `Autopilot: paused (wake expired; resume: user reauthorizes a wake)` and
58
+ notify under the recorded Notification policy when user action is needed.
59
+
60
+ ## Release and install, when applicable
61
+
62
+ Detect applicability once at Align or spec time. Record `Release: <AGENTS.md
63
+ file:line + tag-triggered workflow path + named install hosts> | not applicable
64
+ (<reason>)`. The predicate is an AGENTS.md release rule naming an existing
65
+ tag-triggered workflow. A partial match is not applicable and its reason is
66
+ noted. Install hosts come only from explicit targets; an absent host list is a
67
+ decision hold, not permission to infer hosts. A missing install host list at
68
+ Align or spec time is a decision hold before release authority is presented.
69
+
70
+ Show the `Release:` line in the spec for human approval at gate 1, or the small
71
+ work Align read-back. Copy that decision to `Authority:` in the run record.
72
+ This authority is per run and never carries over to another run or repository.
73
+ The small-work Align read-back names the existing Release and host-mutation
74
+ authority and explicit hosts; silence cannot fill a missing authority or target.
75
+
76
+ After all required feature PRs merge, open one release PR. Default to a patch
77
+ version, or minor if a `feat` commit landed since the last tag. This normal run
78
+ PR gets authored review and diligence of its body against merged PRs, reaches
79
+ merge-ready, then waits for human merge. Once the forge confirms that merge,
80
+ tag and wait for the staged publish. Human npm stage approval is a decision
81
+ hold: agents never run `npm stage approve`. A wake verifies the registry reports
82
+ the expected package and version. Install on the named hosts, verify version
83
+ and roles, then run Close-out last with release and install receipts and the
84
+ installed version.
85
+
86
+ An existing version or tag, failed publish, pending approval, uncertain
87
+ registry result, missing host access, or failed install verification is a
88
+ resumable hold, never success. Tagging, publishing, installation, and host
89
+ mutation require the recorded per-run authority and their existing checks.
90
+
91
+ ## Resume, cancel, and notify
92
+
93
+ At every entry (user message, wake, compaction, or new chat), reconcile the
94
+ owner, authoritative Dispatch, approved revision, PR membership, uncertain
95
+ tags, wakes, publications, and completed receipts under lifecycle and
96
+ run-record before advancing. Only the original driver advances. Wakes do not
97
+ reset attempt budgets and do not grant approvals. Cancel sets `Autopilot: off`,
98
+ stops new actions, and ends the watch under watch §6 with guarded settlement.
99
+ Cancellation does not cancel a running author Dispatch by inference; let it
100
+ report, then settle that exact Dispatch under lifecycle guards without new
101
+ publication.
102
+
103
+ Use the run's recorded Notification policy through `axstack-relay`.
104
+ Decision holds, including spec and npm approval, are always eligible. Across
105
+ implementation and release, merge-ready and merged notifications together are
106
+ capped at two per run; deduplicate by purpose and revision. Healthy ticks stay
107
+ quiet. A failed or uncertain delivery preserves the underlying hold. A relay
108
+ message is only a notification, never authority to approve, merge, or publish.
@@ -6,6 +6,11 @@ Within recorded PR-scoped publication authority, the owner reconciles that
6
6
  receipt against the actual local candidate SHA and base. The owner does not edit
7
7
  the author's candidate; required code changes return to the author.
8
8
 
9
+ Before publication, dispatch `axstack-diligence` under
10
+ [Diligence](diligence.md) to check the author receipt against its evidence
11
+ folder: red/green logs exist, and counts, SHAs, and paths match. Resolve
12
+ `FINDINGS` with the same author before publishing.
13
+
9
14
  Publish the existing commits through `gh stack`. Prefer a fast-forward push.
10
15
  Before a history rewrite, confirm the expected-old remote SHA and use lease
11
16
  protection; a mismatch holds publication. If the push outcome is ambiguous,
@@ -40,3 +45,7 @@ as a `git clone` into a temp directory followed by `orca repo add`; each
40
45
  the directory is gone. Release preparation uses a `release/<version>` worktree
41
46
  of the same registered repo the same way. Release the checkout with
42
47
  `ORCA worktree rm` after its receipt is recorded.
48
+
49
+ For a release PR, dispatch `axstack-diligence` under
50
+ [Diligence](diligence.md) to check the release PR body
51
+ against the merged PRs before publication.
@@ -40,7 +40,11 @@ Validate the configured provider and model at actual launch. If it is
40
40
  unavailable or exhausted, pause affected work, record the gap, and ask the
41
41
  user. Never infer a route from quota state or subscription entitlement. Every
42
42
  substitution requires the user's decision: configured alternatives and native
43
- fallback prose are not defaults.
43
+ fallback prose are not defaults. The only within-class exception is explicit
44
+ model rejection before the first turn: Codex may retry with `--retry-of` using
45
+ the next eligible version in the same class, provider, and effort, recording
46
+ the failed ID, error, and fallback ID. Claude rejection holds. Timeout, quota,
47
+ and auth failures hold.
44
48
 
45
49
  ## Driver and adviser split
46
50
 
@@ -0,0 +1,23 @@
1
+ # Diligence
2
+
3
+ Dispatch `axstack-diligence` through Orca with a pinned brief and evidence paths.
4
+ It is read-only, never authors or edits, and returns `PASS` or `FINDINGS`
5
+ with locations, observed evidence, and limits. A stale or missing receipt is
6
+ not a pass. Keep its first pass independent of other reviewers and workers.
7
+
8
+ For a PR, compare every changed line with the accepted intent and exclusions:
9
+ is it intended and in scope? Check that no contract, rule, or obligation was
10
+ silently weakened or dropped by rewording. Compare the PR body, commit messages,
11
+ and author receipt with the diff: numbers, IDs, versions, test counts, sizes,
12
+ paths, and stale references. Bind the result to the exact head and base.
13
+
14
+ For research, reopen cited sources for answer-changing claims before the
15
+ driver folds verified claims. For a draft spec, compare it with Align decisions
16
+ before user approval: flag anything dropped, added, or softened. For tickets,
17
+ map every spec acceptance item to a capability's acceptance. Before candidate
18
+ publication, compare the author receipt with its evidence folder: red/green
19
+ logs exist, and counts, SHAs, and paths match. For release preparation, compare
20
+ the release PR body with the merged PRs.
21
+
22
+ `FINDINGS` identifies a mismatch for the driver to resolve at the owning phase;
23
+ it does not edit the artifact or create another review round by itself.
@@ -107,8 +107,8 @@ Tracking grants no merge, release, model-substitution, or scope authority.
107
107
 
108
108
  The default 24-hour deadline covers standalone task-owned timers. Stop them at
109
109
  deadline and preserve remaining work; the review automation has no task-owned
110
- deadline. A PR is merge-ready only with the applicable review receipt(s) at
111
- its exact head; green CI or tests alone never make it merge-ready. Merge-ready
110
+ deadline. Merge-ready requires applicable review receipt(s) and current diligence
111
+ `PASS` at the exact head; CI/tests alone are insufficient. Merge-ready
112
112
  differs from merged; human merges.
113
113
 
114
114
  ## Review automation health
@@ -38,23 +38,42 @@ Read `roles.json` from the installed shared root `skills/axstack/`. The installe
38
38
  shape is `{ "version": 1, "preset": "<name>", "roles": [...] }`. Bundled
39
39
  profiles are setup inputs shaped as
40
40
  `{ "version": 1, "roles": [...] }`. A new run records the selected preset and
41
- all 28 role rows once. An active run keeps the exact snapshot until the user
42
- explicitly changes it.
43
-
44
- Select the requested role by stable ID. A missing or null model holds only that role;
45
- never launch a provider default. Launch-by-agent-id routes for which Orca exposes no
41
+ all 32 role rows once. For each role record class, resolved exact ID, source,
42
+ and time. An active run keeps the exact snapshot; resume reuses it without
43
+ re-resolution until the user explicitly changes it.
44
+
45
+ Select the requested role by stable ID. A missing class and missing or null
46
+ model holds only that role; never launch a provider default. Resolve Codex
47
+ classes with `scripts/resolve-models.js`, passing the catalog path explicitly;
48
+ missing or malformed catalogs hold. The first launch of each Claude class uses
49
+ its alias. Read the exact ID from the first assistant turn's `message.model` in
50
+ that worker's own session transcript at
51
+ `~/.claude/projects/<worktree-path-slug>/*.jsonl`; the worktree path slug
52
+ replaces each non-alphanumeric character with `-`. Identify the file by the
53
+ worker's session ID, or use the newest file created after launch. Later launches
54
+ of that class use the recorded exact ID. Before read-back record `alias,
55
+ unresolved`; record an unknown read-back as unknown and hold
56
+ provenance-dependent work. A worker self-report is a labeled last
57
+ resort. Launch-by-agent-id routes for which Orca exposes no
46
58
  `--model` override (today: `grok`, `antigravity`) record `model: null` with an explicit note and are
47
59
  launchable; the run record snapshots the model the TUI reports. Validate provider, model, and effort
48
60
  against the guide and actual launch capability. Stored `modeId` and other
49
61
  permission fields are conservative intent, not proof of effective permission
50
62
  parity or a security boundary. Requested settings, input acceptance, effective
51
63
  settings, and completed work are separate evidence. An unsupported or
52
- unavailable value holds affected work for the user's decision without fallback.
64
+ unavailable value holds affected work for the user's decision except the narrow
65
+ retry below.
53
66
  The single-provider preset's null adviser and round-2 seat are intentional installation data, not
54
67
  readiness failure; because Align and Spec require both adviser receipts, either
55
68
  null adviser still holds those phases. The current chat is the driver and has
56
69
  no role row in any preset.
57
70
 
71
+ Only explicit model rejection before the first turn permits a Codex
72
+ `--retry-of` with the next eligible ID in the same class, provider, and effort.
73
+ Fence the rejected Dispatch and record tried ID, error, and fallback ID in the
74
+ snapshot and reply. Timeout, quota, auth, and other failures hold; Claude
75
+ rejection holds. Apply this to every role, including advisers and judges.
76
+
58
77
  ## Materialize checkouts as worktrees of the registered repo
59
78
 
60
79
  Every reviewer, release, or worker checkout is `ORCA worktree create --repo
@@ -0,0 +1,37 @@
1
+ # Role roster
2
+
3
+ - Chat drives (no role ID); `axstack-owner` owns one PR and
4
+ `axstack-author` its sole writer.
5
+ - `axstack-reviewer-primary` and `axstack-reviewer-secondary` are the ordered
6
+ peer pair. Peer review uses both; authored review uses this table:
7
+
8
+ | Preset | Author class | Reviewer (class/effort) |
9
+ | --- | --- | --- |
10
+ | `mixed` | `codex/sol` | `axstack-reviewer-secondary` (`claude/opus` medium) |
11
+ | `mixed` | `claude/opus` | `axstack-reviewer-primary` (`codex/sol` high) |
12
+ | `codex-only` | `codex/sol` | `axstack-reviewer-secondary` (`codex/luna` xhigh) |
13
+ | `claude-only` | `claude/opus` | `axstack-reviewer-secondary` (`claude/sonnet` high) |
14
+ - `axstack-advisor-astra`/`axstack-advisor-opus` advise and author candidates;
15
+ `axstack-arena-candidate-grok`/
16
+ `axstack-arena-candidate-antigravity` add families.
17
+ `axstack-arena-judge-opus` judges round 1; `axstack-escalation-fable`/`axstack-arena-judge-astra` judge round 2.
18
+ High-stakes/trigger: fresh [contract](contracts.md) session.
19
+ `axstack-auditor` audits; `axstack-checker` reports discrepancies.
20
+ - `axstack-explainer`/`axstack-explainer-review`: explain/review.
21
+ - `axstack-diligence`: read-only [diligence checks](diligence.md) for every PR
22
+ review round and bounded research, spec, ticket, receipt, and release claims.
23
+ - `axstack-ui-verifier`: [UI checks](ui-verification.md).
24
+ - `axstack-auditor`/`axstack-research-requirements`/
25
+ `axstack-research-code`/`axstack-research-web`/
26
+ `axstack-explore-execution`/`axstack-monitor`:
27
+ `claude/sonnet` high in mixed/claude-only.
28
+ `axstack-monitor`: standalone watch never sends; chat-run watch: bounded
29
+ internal reports to its Run and original driver.
30
+ - Sol pairs `axstack-auditor-sol`/`axstack-research-code-sol`/
31
+ `axstack-explore-execution-sol`: `codex/sol` high in
32
+ mixed/codex-only; intentionally absent in claude-only. Dispatch each
33
+ independently from its Sonnet seat on the same bounded brief without
34
+ cross-reading. The driver reconciles findings per claim, never averages.
35
+ Record intentional absence and continue with Sonnet alone; a configured
36
+ but unavailable seat holds only its affected work.
37
+ - `axstack-debug-investigator-1..4` probe L1 briefs.
@@ -11,45 +11,33 @@ skills root, or an explicit user selection in the run record. Missing or contrad
11
11
  a setup gap: hold. Never infer from live profiles or `list_profiles`, harness,
12
12
  tools, credentials, quota, subscription, or default to `mixed`.
13
13
 
14
- At start, snapshot all 28 role IDs with provider/model/mode/effort; absent
14
+ At start, snapshot all 32 role IDs with provider/modelClass/model/mode/effort; absent
15
15
  or unconfigured roles are recorded explicitly; never default.
16
16
  Such a role holds only its work. Later installed or changed roles need an
17
17
  explicit user decision to enter the snapshot. Live profiles
18
18
  are authoritative at snapshot time and for availability; bundled presets are setup
19
19
  inputs, not runtime proof.
20
+ For each role record class, resolved exact ID, source (catalog, transcript, or
21
+ pin), and time. Codex classes resolve through
22
+ `skills/axstack/scripts/resolve-models.js` with an explicit catalog
23
+ path; missing or malformed catalog holds. Claude classes start as `alias,
24
+ unresolved` until transcript read-back. Resume must reuse the snapshot and
25
+ never re-resolve it.
20
26
 
21
27
  Preset changes apply to new runs only; an active run keeps its snapshot.
22
28
  Changing it or replacing a session needs an explicit user decision and
23
29
  revalidation. Unavailable models, efforts, roles, or overrides hold only affected
24
30
  work; no automatic fallback, quota routing, subscription inference, or silent
25
- provider/model/effort substitution.
26
-
27
- - Chat drives (no role ID); `axstack-owner` owns one PR and
28
- `axstack-author` its sole writer.
29
- - `axstack-reviewer-primary` and `axstack-reviewer-secondary` are the ordered
30
- peer pair. Peer review uses both; authored review uses this table:
31
-
32
- | Preset | Author | Reviewer (model/effort) |
33
- | --- | --- | --- |
34
- | `mixed` | Codex / Sol (`codex/gpt-6-sol`) | `axstack-reviewer-secondary` (`claude/claude-opus-5-5` medium) |
35
- | `mixed` | Claude / Opus (`claude/claude-opus-5-5`) | `axstack-reviewer-primary` (`codex/gpt-6-sol` high) |
36
- | `codex-only` | Codex / Sol (`codex/gpt-6-sol`) | `axstack-reviewer-secondary` (`codex/gpt-6-luna` xhigh) |
37
- | `claude-only` | Claude / Opus (`claude/claude-opus-5-5`) | `axstack-reviewer-secondary` (`claude/claude-sonnet-5-5` high) |
38
- - `axstack-advisor-astra`/`axstack-advisor-opus` advise and author candidates;
39
- `axstack-arena-candidate-grok`/
40
- `axstack-arena-candidate-antigravity` add families.
41
- `axstack-arena-judge-opus` judges round 1; `axstack-escalation-fable`/`axstack-arena-judge-astra` judge round 2.
42
- High-stakes/trigger: fresh [contract](contracts.md) session.
43
- `axstack-auditor` audits; `axstack-checker` reports discrepancies.
44
- - `axstack-explainer`/`axstack-explainer-review`: explain/review.
45
- - `axstack-ui-verifier`: [UI checks](ui-verification.md).
46
- - `axstack-research-requirements`/`axstack-research-web`/`axstack-monitor`:
47
- Sonnet 5.5 high in mixed/claude-only.
48
- `axstack-monitor`: standalone watch never sends; chat-run watch: bounded
49
- internal reports to its Run and original driver.
50
- - `axstack-debug-investigator-1..4` probe L1 briefs.
51
-
52
- Provenance is matched on provider/model ID; effort never maps. Missing table-row
31
+ provider/model/effort substitution. Only
32
+ explicit model rejection before the first turn permits Codex `--retry-of` with
33
+ the next eligible ID in the same class, provider, and effort. Fence the failed
34
+ Dispatch and record tried ID, error, and fallback ID in the snapshot and reply.
35
+ Timeout, quota, auth, and other failures hold; Claude rejection holds.
36
+
37
+ Load the [Role roster](role-roster.md) for configured roles and authored-review pairings.
38
+
39
+ Provenance is matched on provider/model class derived from the recorded exact ID;
40
+ effort never maps. Missing table-row
53
41
  provenance is unsupported and `INCOMPLETE`; report it and ask the user. Never
54
42
  infer from slot, driver, owner, or provider. Author and owner never review.
55
43
 
@@ -98,7 +86,7 @@ reason in the run record, or in the brief for tiny direct work.
98
86
  Require an approved spec plus a ticket map tied to that exact spec
99
87
  revision, with acceptance checks and dependencies in the explicitly selected
100
88
  Markdown, GitHub Issues, or Linear store. Prepare via `axstack-align` -> `axstack-spec`
101
- (one approval) -> `axstack-tickets` -> handoff, then stop.
89
+ (one approval) -> `axstack-tickets` -> handoff, then continue under autopilot when eligible.
102
90
  - **Small:** clear, bounded one-PR work. The driver captures the named
103
91
  **small-change intent** from the current request or user-chosen existing
104
92
  issue plus explicit acceptance checks and exclusions, snapshots it once, and
@@ -119,7 +107,7 @@ not alone a formal spec trigger. Hold affected unsafe work while reassessing.
119
107
  ## Lifecycle routes (mode-specific scope identity required)
120
108
 
121
109
  - Preparation: substantial work follows the align -> spec -> tickets ->
122
- handoff path above, then stops; small work uses the driver-captured
110
+ handoff path above, then continues under autopilot when eligible; small work uses the driver-captured
123
111
  small-change intent.
124
112
  - Execution: with its identity present, `axstack-implement` ->
125
113
  `axstack-review` -> `axstack-watch`.
@@ -151,6 +151,8 @@ Authority: <who authorized which mutation>
151
151
  Intent: <approved spec rev | small-change intent | adopted snapshot | peer/read-only mode>
152
152
  Routing: <preset + source + snapshot ref>
153
153
  Notification policy: <none | transport/target label/host/instructions path>
154
+ Autopilot: on | paused (<hold>; resume: <condition>) | off (cancelled <ts>); next: <step>
155
+ Release: <AGENTS.md file:line + tag-triggered workflow path + named install hosts> | not applicable (<reason>)
154
156
  Source base: <exact revision or source identity>
155
157
  IDs: <repo/project + workspace/agent receipt pointers>
156
158
  Worktrees in other repositories: <per-run repository and worktree IDs or none>
@@ -0,0 +1,46 @@
1
+ import { readFileSync } from 'node:fs';
2
+
3
+ const args = process.argv.slice(2);
4
+ const option = (name) => {
5
+ const index = args.indexOf(name);
6
+ return index < 0 ? null : args[index + 1];
7
+ };
8
+ const path = option('--catalog');
9
+ const modelClass = option('--class');
10
+ const effort = option('--effort');
11
+
12
+ try {
13
+ if (!path || !['astra', 'sol', 'luna'].includes(modelClass) || !effort) {
14
+ throw new Error('expected --catalog path --class astra|sol|luna --effort level');
15
+ }
16
+ const catalog = JSON.parse(readFileSync(path, 'utf8'));
17
+ if (!Array.isArray(catalog.models) || typeof catalog.client_version !== 'string'
18
+ || typeof catalog.fetched_at !== 'string') {
19
+ throw new Error('malformed catalog');
20
+ }
21
+ const classPattern = new RegExp(`^gpt-(\\d+(?:\\.\\d+)*)-${modelClass}$`);
22
+ const candidates = catalog.models
23
+ .filter((entry) => entry && entry.visibility === 'list'
24
+ && typeof entry.slug === 'string'
25
+ && classPattern.test(entry.slug)
26
+ && Array.isArray(entry.supported_reasoning_levels)
27
+ && entry.supported_reasoning_levels.some((level) => level?.effort === effort))
28
+ .map((entry) => entry.slug)
29
+ .sort((left, right) => {
30
+ const a = left.slice(4, -(modelClass.length + 1)).split('.').map(Number);
31
+ const b = right.slice(4, -(modelClass.length + 1)).split('.').map(Number);
32
+ for (let i = 0; i < Math.max(a.length, b.length); i++) {
33
+ const difference = (b[i] ?? 0) - (a[i] ?? 0);
34
+ if (difference) return difference;
35
+ }
36
+ return 0;
37
+ });
38
+ if (!candidates.length) throw new Error(`no eligible ${modelClass} model in catalog`);
39
+ console.log(JSON.stringify({
40
+ provider: 'codex', modelClass, effort, model: candidates[0], candidates,
41
+ catalog: { path, client_version: catalog.client_version, fetched_at: catalog.fetched_at },
42
+ }));
43
+ } catch (error) {
44
+ console.error(`model catalog resolution hold: ${error.message}`);
45
+ process.exitCode = 1;
46
+ }
@@ -5,6 +5,9 @@ description: When exploring or planning engineering work, use axstack-align to s
5
5
 
6
6
  # Align
7
7
 
8
+ For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
9
+ for phase continuation and holds.
10
+
8
11
  On driver entry, sweep under [Workspace hygiene](../axstack/references/workspace-hygiene.md); dispatched workers do not sweep.
9
12
  For every dispatch brief, name its private `<run dir>/evidence/<dispatch>/` folder.
10
13
 
@@ -166,7 +169,7 @@ record or spec. Read-only scope keeps proposed documentation in the permitted
166
169
  private record or response. Documentation is neither implementation nor spec
167
170
  approval; record chosen document names and paths once per run.
168
171
 
169
- ## Read back, classify, and stop
172
+ ## Read back, classify, and route
170
173
 
171
174
  1. Read back the decisions, constraints, exclusions, and remaining evidence
172
175
  gaps. For substantial work, this summary becomes part of the draft spec in
@@ -189,6 +192,7 @@ approval; record chosen document names and paths once per run.
189
192
  [Orca runtime](../axstack/references/orca-runtime.md) immediately before
190
193
  actual dispatch. Alignment completion never dispatches a recipient.
191
194
 
192
- Alignment stops for both sizes only when the handoff is usable, its next scope
193
- identity is explicit, and execution has not started. The user invokes
194
- `axstack-implement` to execute.
195
+ Alignment completes for both sizes only when the handoff is usable and its next
196
+ scope identity is explicit. An eligible delivery run continues under Autopilot;
197
+ an explicit stop-after-Align request ends here. Substantial work continues to
198
+ Spec, and small work continues from its small-change intent to Implement.
@@ -28,8 +28,13 @@ immediately before an actual auditor profile or session dispatch. Ordinary
28
28
  audit reading and record writing do not load it, and the auditor never
29
29
  dispatches.
30
30
 
31
- Core owns the `axstack-auditor` profile (codex/gpt-6-luna xhigh) and its
32
- invocation. This skill governs what that auditor reads, measures, and proposes.
31
+ Core owns the `axstack-auditor` profile (claude/sonnet high in
32
+ mixed/claude-only; codex/luna xhigh in codex-only) and its invocation.
33
+ This skill governs what that auditor reads, measures, and proposes.
34
+ Dispatch `axstack-auditor` and `axstack-auditor-sol` independently on the same
35
+ bounded brief, without cross-reading. The driver reconciles findings per claim;
36
+ never average verdicts. Record an intentionally absent Sol seat and continue
37
+ with the base auditor alone; a configured but unavailable seat holds its work.
33
38
  The user-chosen improvement mode is a tested, independently reviewed PR that a
34
39
  human merges.
35
40
 
@@ -86,6 +91,14 @@ counts with denominators plus the evidence behind the count:
86
91
  evidence path is absent, or `UNKNOWN` with the reason when its records are
87
92
  unavailable.
88
93
  - Independent exact-revision review status and unresolved findings.
94
+ For authored PRs, measure the selected reviewer against this class pairing:
95
+
96
+ | Preset | Author class | Reviewer (class/effort) |
97
+ | --- | --- | --- |
98
+ | `mixed` | `codex/sol` | `axstack-reviewer-secondary` (`claude/opus` medium) |
99
+ | `mixed` | `claude/opus` | `axstack-reviewer-primary` (`codex/sol` high) |
100
+ | `codex-only` | `codex/sol` | `axstack-reviewer-secondary` (`codex/luna` xhigh) |
101
+ | `claude-only` | `claude/opus` | `axstack-reviewer-secondary` (`claude/sonnet` high) |
89
102
  - Simplification applicability determinations evidenced / total candidate diffs,
90
103
  and complete simplification receipts / total candidates, broken down as
91
104
  `applied`, `not-applicable`, or `UNKNOWN` with the reason. This measures
@@ -5,6 +5,9 @@ description: When an approved task is ready to build or repair, use axstack-impl
5
5
 
6
6
  # Implement
7
7
 
8
+ For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
9
+ for phase continuation and holds.
10
+
8
11
  On driver entry, sweep under [Workspace hygiene](../axstack/references/workspace-hygiene.md); dispatched workers do not sweep.
9
12
  For every dispatch brief, name its private `<run dir>/evidence/<dispatch>/` folder.
10
13
  Include [Safe deletion](../axstack/references/workspace-hygiene.md#safe-deletion) in author briefs.
@@ -206,27 +209,43 @@ For each PR:
206
209
  `REQUEST_CHANGES`, a failed required check, or post-readiness feedback returns
207
210
  findings to the same author for a new revision, increments `repairs`, and
208
211
  returns to step 1. `INCOMPLETE`, a provenance gap, unavailable model, serious
209
- risk, or the third `REQUEST_CHANGES` on one PR records `held`. A changed
210
- parent sends its child back to step 1.
212
+ risk, or the third review round with `REQUEST_CHANGES` and/or diligence
213
+ `FINDINGS` on one PR records `held`. A changed parent sends its child back
214
+ to step 1.
215
+ Merge-ready also requires a current diligence `PASS` at that head; diligence
216
+ `FINDINGS` return to the same author within the review round.
217
+ A round with reviewer `REQUEST_CHANGES` and/or diligence `FINDINGS` increments
218
+ `repairs` once and counts once toward the third-round hold.
211
219
 
212
220
  One run-level completion wait covers every unsettled Dispatch; the bounded
213
- forge check wait is the only other wait. End a turn only when every required PR
221
+ forge check wait is the only other implementation wait. The eligible run arms
222
+ one maintain-mode chat-run watch at its first published PR; that watch owns its
223
+ 10-minute harness wake or Orca fallback. End a turn only when every required PR
214
224
  is `merge-ready` or `held`. Under the recorded Notification policy,
215
- `axstack-relay` sends only a serious risk immediately or a genuine blocked
216
- operation that needs user intervention after bounded safe recovery. Questions,
217
- spec approvals, progress, CI pending, merge-ready, merged, and completion stay
225
+ `axstack-relay` sends only a serious risk immediately, a genuine blocked
226
+ operation needing user intervention after bounded safe recovery, or the
227
+ decision holds and capped milestones named by the recorded Notification policy.
228
+ Routine questions stay in Orca. Progress, CI pending, and completion always stay
218
229
  in Orca.
230
+ Only the bounded categories—user-decision holds (including spec approval),
231
+ serious-risk holds, and at most two merge-ready/merged milestones per run—may
232
+ be relayed under the recorded Notification policy.
219
233
 
220
234
  Merge-ready is the human boundary: the user merges, bottom-up for a stack. The
221
- driver resumes on the user's next message or `/axstack-watch`; no Orca merge
222
- wake exists today. Re-read forge state: record forge-merged PRs as `merged`;
235
+ driver resumes on the user's next message, `/axstack-watch`, or the armed
236
+ chat-run watch wake; no Orca merge wake exists today.
237
+ Re-read forge state: record forge-merged PRs as `merged`;
223
238
  changed heads or feedback return to step 1; retain useful author work before Close-out.
224
- Run Close-out once only after every required PR is forge-merged and acceptance
225
- passes. It settles workers, records counts, makes the auditor decision and
239
+ Run Close-out once only after every required PR is forge-merged, the run's
240
+ Release step is settled or not applicable, and acceptance passes. It settles
241
+ workers, records counts, makes the auditor decision and
226
242
  settlement, releases worktrees, closes eligible tickets, and archives the run.
243
+ Without an Autopilot or Release record, the Release step is not applicable for
244
+ both Close-out and run completion.
227
245
 
228
246
  The loop requires the `mixed` two-provider authored-review row. `codex-only` or
229
247
  `claude-only` holds at step (3) for an explicit user routing choice, with no
230
248
  substitution or same-provider review. Derived PR states are `authoring |
231
249
  published | in-review | repairing(n) | merge-ready | merged | held`. The run is
232
- done only when every required PR is forge-merged and Close-out has receipts.
250
+ done only when every required PR is forge-merged, the run's Release step is
251
+ settled or not applicable, and Close-out has receipts.
@@ -19,6 +19,11 @@ role dispatch, load the [Orca runtime
19
19
  sequence](../axstack/references/orca-runtime.md). Use existing
20
20
  `axstack-explore-codebase` or `axstack-research-code` roles only when their
21
21
  specialization materially helps; create no new profile.
22
+ When dispatching `axstack-research-code`, dispatch `axstack-research-code-sol`
23
+ independently on the same bounded brief without cross-reading. The driver
24
+ reconciles findings per claim and never averages them. Record an intentionally
25
+ absent Sol pair and proceed with the base seat alone; a configured but
26
+ unavailable pair holds its work.
22
27
 
23
28
  ## 1. Bound discovery
24
29
 
@@ -5,6 +5,9 @@ description: When the user requests a relay message or test, or an authorized no
5
5
 
6
6
  # Relay
7
7
 
8
+ For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
9
+ for phase continuation and holds.
10
+
8
11
  Send normal messages, transport tests, and authorized notifications to the
9
12
  user through Hermes' native one-way `hermes send`. This is an inline caller
10
13
  procedure: it creates no driver, team, owner, auditor, monitor, child session,
@@ -25,8 +28,12 @@ Choose the applicable message type:
25
28
  in the caller's private notification policy. State the issue, impact, and the
26
29
  answer or action needed.
27
30
  - **Routine run events:** questions, spec approvals, progress, CI pending,
28
- merge-ready, merged, and completion stay in Orca. They never become proactive
29
- relay messages merely because the run is waiting.
31
+ merge-ready, merged, and completion stay in Orca unless the recorded
32
+ Notification policy names it. A policy may name only user-decision holds and
33
+ at most two merge-ready/merged milestones per run; deduplicate across implementation
34
+ and release. Progress, CI pending, and completion are never eligible merely
35
+ because a policy exists. They never become proactive relay messages merely
36
+ because the run is waiting.
30
37
 
31
38
  Verify the transport, execution host, and intended recipient from the user's
32
39
  request, trusted caller context, or an existing private notification policy.
@@ -29,8 +29,9 @@ is part of research.
29
29
 
30
30
  2. **Fan out research:** A single factual lookup stays in the current chat.
31
31
  Every other research run dispatches every configured research branch through
32
- Orca: requirements (Claude), code (Codex), web (Claude), web-google
33
- (Gemini/Antigravity, with Google Search built in), and X (Grok).
32
+ Orca: requirements, code, and web (Sonnet high in mixed/claude-only;
33
+ Codex in codex-only), web-google (Gemini/Antigravity, with Google Search
34
+ built in), and X (Grok).
34
35
  Give each branch one owner, allow no cross-reading, and require a cited note
35
36
  with a URL and access date per claim; re-open sources and never trust a search
36
37
  summary. The driver reconciles agreements/disagreements per claim.
@@ -41,16 +42,27 @@ is part of research.
41
42
 
42
43
  - `axstack-research-requirements`: requirements and intent.
43
44
  - `axstack-research-code`: code behavior.
45
+ - `axstack-research-code-sol`: independent Sol code investigation.
44
46
  - `axstack-research-web`: web and external sources.
45
47
  - `axstack-research-web-google`: Google-Search-grounded web sources via Gemini/Antigravity.
46
48
  - `axstack-research-x`: X (Twitter) posts and threads via Grok; cite post URLs and dates.
47
49
  - `axstack-explore-codebase`: broad codebase mapping.
48
50
  - `axstack-explore-execution`: execution and runtime traces.
51
+ - `axstack-explore-execution-sol`: independent Sol execution investigation.
52
+
53
+ When dispatching `axstack-research-code` or `axstack-explore-execution`,
54
+ dispatch its `-sol` pair independently on the same bounded brief without
55
+ cross-reading. The driver reconciles agreement and disagreement per claim,
56
+ never averaging findings. Record an intentionally absent pair and proceed
57
+ with the base seat alone; a configured but unavailable seat holds its work.
49
58
 
50
59
  3. **Gather primary source evidence.** Inspect the actual documentation, code,
51
60
  or tool output for every answer-changing claim. Apply the source standards
52
61
  for citations, freshness, revisions, and access dates. Continue until each
53
62
  material claim has direct evidence or a named evidence gap.
63
+ Before the driver folds verified claims, dispatch `axstack-diligence` under
64
+ [Diligence](../axstack/references/diligence.md) to reopen cited sources for
65
+ answer-changing claims and flag mismatches.
54
66
 
55
67
  4. **Form the verdict.** Mark every material claim as **verified**,
56
68
  **inference**, or **unverified** using the source standards. Derive