axstack 0.20.30 → 0.20.31
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +2 -2
- package/bin/axstack.js +1 -0
- package/docs/installation.md +12 -6
- package/docs/workflows.md +29 -10
- package/package.json +1 -1
- package/profiles/presets/claude-only.json +46 -46
- package/profiles/presets/codex-only.json +50 -50
- package/profiles/presets/mixed.json +54 -54
- package/skills/axstack/references/autopilot.md +108 -0
- package/skills/axstack/references/contracts.md +5 -1
- package/skills/axstack/references/orca-runtime.md +25 -6
- package/skills/axstack/references/role-roster.md +7 -7
- package/skills/axstack/references/routing.md +16 -5
- package/skills/axstack/references/run-record.md +2 -0
- package/skills/axstack/scripts/resolve-models.js +46 -0
- package/skills/axstack-align/SKILL.md +8 -4
- package/skills/axstack-audit/SKILL.md +10 -2
- package/skills/axstack-implement/SKILL.md +23 -9
- package/skills/axstack-relay/SKILL.md +9 -2
- package/skills/axstack-review/SKILL.md +13 -7
- package/skills/axstack-spec/SKILL.md +5 -1
- package/skills/axstack-tickets/SKILL.md +6 -3
- package/skills/axstack-watch/SKILL.md +26 -13
- package/skills/axstack-watch/references/watch-runtime.md +11 -4
- package/src/installer.js +8 -0
- package/src/roles.js +38 -10
|
@@ -0,0 +1,108 @@
|
|
|
1
|
+
# Autopilot
|
|
2
|
+
|
|
3
|
+
This reference applies to authorized engineering-delivery runs. The original
|
|
4
|
+
driver remains the sole run-record writer and phase router in the same chat;
|
|
5
|
+
phase completion is not a native ownership handoff. Explicit planning-only,
|
|
6
|
+
read-only, stop-after-phase, observation-only, and peer requests retain their
|
|
7
|
+
selected boundary. A status question such as "what's left" is observation,
|
|
8
|
+
not a mode change.
|
|
9
|
+
|
|
10
|
+
## Advance and hold
|
|
11
|
+
|
|
12
|
+
Advance only after the finishing phase returns its completed identity (a
|
|
13
|
+
small-change intent, approved spec, matching ticket map, merge-ready or merged
|
|
14
|
+
state) and the run record has no open hold. A hold from any phase stops the run:
|
|
15
|
+
record its reason, owner, and resume condition, then take no dependent action.
|
|
16
|
+
That covers tracker access, adviser or arena-seat availability, diligence
|
|
17
|
+
FINDINGS when the phase records a hold, CI-wait timeout, readiness UNKNOWN,
|
|
18
|
+
dismissed approval, wake or cleanup uncertainty, single-provider routing, an
|
|
19
|
+
existing tag or version, and failed publish. Diligence FINDINGS during implement
|
|
20
|
+
follow its §6 repair route; at spec, tickets, or release preparation the driver
|
|
21
|
+
resolves them before advancing, and only a recorded hold pauses autopilot.
|
|
22
|
+
|
|
23
|
+
Record `Autopilot: on | paused (<hold>; resume: <condition>) | off (cancelled
|
|
24
|
+
<ts>)` and the next step in the private run record. A user answer to the hold
|
|
25
|
+
resumes after reconciliation; silence does not.
|
|
26
|
+
Awaiting human spec approval records `Autopilot: paused (spec approval; resume:
|
|
27
|
+
human approval)` as a decision hold eligible under the Notification policy.
|
|
28
|
+
|
|
29
|
+
## Phase sequence
|
|
30
|
+
|
|
31
|
+
- Small: Align read-back, small-change intent, implement, watch in maintain
|
|
32
|
+
mode, human merge. An opted-in Align refinement is part of read-back.
|
|
33
|
+
- Substantial: Align, spec draft with advisers and diligence, human spec
|
|
34
|
+
approval at gate 1, tickets with diligence, implement, watch in maintain mode,
|
|
35
|
+
merge-ready, human merge. An opted-in Align refinement is part of gate 1.
|
|
36
|
+
|
|
37
|
+
Do not seek another phase-start instruction after a completed identity.
|
|
38
|
+
Spec approval is always the human's decision. Every PR merge is the human's,
|
|
39
|
+
including a release PR and each PR in a stack, bottom-up.
|
|
40
|
+
|
|
41
|
+
## Implement into maintain watch
|
|
42
|
+
|
|
43
|
+
When implement publishes the run's first PR, arm exactly one `axstack-watch`
|
|
44
|
+
chat-run in authorized maintain mode. Use the 10-minute harness wake, with the
|
|
45
|
+
existing Orca fallback when unavailable. Later run PRs join after verified
|
|
46
|
+
publication readback; an explicitly adopted PR joins only with its maintenance
|
|
47
|
+
snapshot. The original driver alone routes work; one author writes each
|
|
48
|
+
candidate. Until a PR is merge-ready, wakes feed implement §6 step 4. After
|
|
49
|
+
merge-ready, watch §5 maintenance repairs feedback, rebases when the base moves,
|
|
50
|
+
keeps CI green, and checks approvals without re-requesting human review.
|
|
51
|
+
|
|
52
|
+
Maintain is the default mode for run-created PRs. End the chat-run watch when
|
|
53
|
+
every watched PR is merged or closed and the run's release step is settled or
|
|
54
|
+
not applicable, or when the user cancels. Expiry is a recorded stop with
|
|
55
|
+
resumable state, never a silent renewal. A required PR closed without merging
|
|
56
|
+
is incomplete scope; it does not make the run release-eligible. On wake expiry
|
|
57
|
+
record `Autopilot: paused (wake expired; resume: user reauthorizes a wake)` and
|
|
58
|
+
notify under the recorded Notification policy when user action is needed.
|
|
59
|
+
|
|
60
|
+
## Release and install, when applicable
|
|
61
|
+
|
|
62
|
+
Detect applicability once at Align or spec time. Record `Release: <AGENTS.md
|
|
63
|
+
file:line + tag-triggered workflow path + named install hosts> | not applicable
|
|
64
|
+
(<reason>)`. The predicate is an AGENTS.md release rule naming an existing
|
|
65
|
+
tag-triggered workflow. A partial match is not applicable and its reason is
|
|
66
|
+
noted. Install hosts come only from explicit targets; an absent host list is a
|
|
67
|
+
decision hold, not permission to infer hosts. A missing install host list at
|
|
68
|
+
Align or spec time is a decision hold before release authority is presented.
|
|
69
|
+
|
|
70
|
+
Show the `Release:` line in the spec for human approval at gate 1, or the small
|
|
71
|
+
work Align read-back. Copy that decision to `Authority:` in the run record.
|
|
72
|
+
This authority is per run and never carries over to another run or repository.
|
|
73
|
+
The small-work Align read-back names the existing Release and host-mutation
|
|
74
|
+
authority and explicit hosts; silence cannot fill a missing authority or target.
|
|
75
|
+
|
|
76
|
+
After all required feature PRs merge, open one release PR. Default to a patch
|
|
77
|
+
version, or minor if a `feat` commit landed since the last tag. This normal run
|
|
78
|
+
PR gets authored review and diligence of its body against merged PRs, reaches
|
|
79
|
+
merge-ready, then waits for human merge. Once the forge confirms that merge,
|
|
80
|
+
tag and wait for the staged publish. Human npm stage approval is a decision
|
|
81
|
+
hold: agents never run `npm stage approve`. A wake verifies the registry reports
|
|
82
|
+
the expected package and version. Install on the named hosts, verify version
|
|
83
|
+
and roles, then run Close-out last with release and install receipts and the
|
|
84
|
+
installed version.
|
|
85
|
+
|
|
86
|
+
An existing version or tag, failed publish, pending approval, uncertain
|
|
87
|
+
registry result, missing host access, or failed install verification is a
|
|
88
|
+
resumable hold, never success. Tagging, publishing, installation, and host
|
|
89
|
+
mutation require the recorded per-run authority and their existing checks.
|
|
90
|
+
|
|
91
|
+
## Resume, cancel, and notify
|
|
92
|
+
|
|
93
|
+
At every entry (user message, wake, compaction, or new chat), reconcile the
|
|
94
|
+
owner, authoritative Dispatch, approved revision, PR membership, uncertain
|
|
95
|
+
tags, wakes, publications, and completed receipts under lifecycle and
|
|
96
|
+
run-record before advancing. Only the original driver advances. Wakes do not
|
|
97
|
+
reset attempt budgets and do not grant approvals. Cancel sets `Autopilot: off`,
|
|
98
|
+
stops new actions, and ends the watch under watch §6 with guarded settlement.
|
|
99
|
+
Cancellation does not cancel a running author Dispatch by inference; let it
|
|
100
|
+
report, then settle that exact Dispatch under lifecycle guards without new
|
|
101
|
+
publication.
|
|
102
|
+
|
|
103
|
+
Use the run's recorded Notification policy through `axstack-relay`.
|
|
104
|
+
Decision holds, including spec and npm approval, are always eligible. Across
|
|
105
|
+
implementation and release, merge-ready and merged notifications together are
|
|
106
|
+
capped at two per run; deduplicate by purpose and revision. Healthy ticks stay
|
|
107
|
+
quiet. A failed or uncertain delivery preserves the underlying hold. A relay
|
|
108
|
+
message is only a notification, never authority to approve, merge, or publish.
|
|
@@ -40,7 +40,11 @@ Validate the configured provider and model at actual launch. If it is
|
|
|
40
40
|
unavailable or exhausted, pause affected work, record the gap, and ask the
|
|
41
41
|
user. Never infer a route from quota state or subscription entitlement. Every
|
|
42
42
|
substitution requires the user's decision: configured alternatives and native
|
|
43
|
-
fallback prose are not defaults.
|
|
43
|
+
fallback prose are not defaults. The only within-class exception is explicit
|
|
44
|
+
model rejection before the first turn: Codex may retry with `--retry-of` using
|
|
45
|
+
the next eligible version in the same class, provider, and effort, recording
|
|
46
|
+
the failed ID, error, and fallback ID. Claude rejection holds. Timeout, quota,
|
|
47
|
+
and auth failures hold.
|
|
44
48
|
|
|
45
49
|
## Driver and adviser split
|
|
46
50
|
|
|
@@ -38,23 +38,42 @@ Read `roles.json` from the installed shared root `skills/axstack/`. The installe
|
|
|
38
38
|
shape is `{ "version": 1, "preset": "<name>", "roles": [...] }`. Bundled
|
|
39
39
|
profiles are setup inputs shaped as
|
|
40
40
|
`{ "version": 1, "roles": [...] }`. A new run records the selected preset and
|
|
41
|
-
all 32 role rows once.
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
41
|
+
all 32 role rows once. For each role record class, resolved exact ID, source,
|
|
42
|
+
and time. An active run keeps the exact snapshot; resume reuses it without
|
|
43
|
+
re-resolution until the user explicitly changes it.
|
|
44
|
+
|
|
45
|
+
Select the requested role by stable ID. A missing class and missing or null
|
|
46
|
+
model holds only that role; never launch a provider default. Resolve Codex
|
|
47
|
+
classes with `scripts/resolve-models.js`, passing the catalog path explicitly;
|
|
48
|
+
missing or malformed catalogs hold. The first launch of each Claude class uses
|
|
49
|
+
its alias. Read the exact ID from the first assistant turn's `message.model` in
|
|
50
|
+
that worker's own session transcript at
|
|
51
|
+
`~/.claude/projects/<worktree-path-slug>/*.jsonl`; the worktree path slug
|
|
52
|
+
replaces each non-alphanumeric character with `-`. Identify the file by the
|
|
53
|
+
worker's session ID, or use the newest file created after launch. Later launches
|
|
54
|
+
of that class use the recorded exact ID. Before read-back record `alias,
|
|
55
|
+
unresolved`; record an unknown read-back as unknown and hold
|
|
56
|
+
provenance-dependent work. A worker self-report is a labeled last
|
|
57
|
+
resort. Launch-by-agent-id routes for which Orca exposes no
|
|
46
58
|
`--model` override (today: `grok`, `antigravity`) record `model: null` with an explicit note and are
|
|
47
59
|
launchable; the run record snapshots the model the TUI reports. Validate provider, model, and effort
|
|
48
60
|
against the guide and actual launch capability. Stored `modeId` and other
|
|
49
61
|
permission fields are conservative intent, not proof of effective permission
|
|
50
62
|
parity or a security boundary. Requested settings, input acceptance, effective
|
|
51
63
|
settings, and completed work are separate evidence. An unsupported or
|
|
52
|
-
unavailable value holds affected work for the user's decision
|
|
64
|
+
unavailable value holds affected work for the user's decision except the narrow
|
|
65
|
+
retry below.
|
|
53
66
|
The single-provider preset's null adviser and round-2 seat are intentional installation data, not
|
|
54
67
|
readiness failure; because Align and Spec require both adviser receipts, either
|
|
55
68
|
null adviser still holds those phases. The current chat is the driver and has
|
|
56
69
|
no role row in any preset.
|
|
57
70
|
|
|
71
|
+
Only explicit model rejection before the first turn permits a Codex
|
|
72
|
+
`--retry-of` with the next eligible ID in the same class, provider, and effort.
|
|
73
|
+
Fence the rejected Dispatch and record tried ID, error, and fallback ID in the
|
|
74
|
+
snapshot and reply. Timeout, quota, auth, and other failures hold; Claude
|
|
75
|
+
rejection holds. Apply this to every role, including advisers and judges.
|
|
76
|
+
|
|
58
77
|
## Materialize checkouts as worktrees of the registered repo
|
|
59
78
|
|
|
60
79
|
Every reviewer, release, or worker checkout is `ORCA worktree create --repo
|
|
@@ -5,12 +5,12 @@
|
|
|
5
5
|
- `axstack-reviewer-primary` and `axstack-reviewer-secondary` are the ordered
|
|
6
6
|
peer pair. Peer review uses both; authored review uses this table:
|
|
7
7
|
|
|
8
|
-
| Preset | Author | Reviewer (
|
|
8
|
+
| Preset | Author class | Reviewer (class/effort) |
|
|
9
9
|
| --- | --- | --- |
|
|
10
|
-
| `mixed` |
|
|
11
|
-
| `mixed` |
|
|
12
|
-
| `codex-only` |
|
|
13
|
-
| `claude-only` |
|
|
10
|
+
| `mixed` | `codex/sol` | `axstack-reviewer-secondary` (`claude/opus` medium) |
|
|
11
|
+
| `mixed` | `claude/opus` | `axstack-reviewer-primary` (`codex/sol` high) |
|
|
12
|
+
| `codex-only` | `codex/sol` | `axstack-reviewer-secondary` (`codex/luna` xhigh) |
|
|
13
|
+
| `claude-only` | `claude/opus` | `axstack-reviewer-secondary` (`claude/sonnet` high) |
|
|
14
14
|
- `axstack-advisor-astra`/`axstack-advisor-opus` advise and author candidates;
|
|
15
15
|
`axstack-arena-candidate-grok`/
|
|
16
16
|
`axstack-arena-candidate-antigravity` add families.
|
|
@@ -24,11 +24,11 @@
|
|
|
24
24
|
- `axstack-auditor`/`axstack-research-requirements`/
|
|
25
25
|
`axstack-research-code`/`axstack-research-web`/
|
|
26
26
|
`axstack-explore-execution`/`axstack-monitor`:
|
|
27
|
-
`claude
|
|
27
|
+
`claude/sonnet` high in mixed/claude-only.
|
|
28
28
|
`axstack-monitor`: standalone watch never sends; chat-run watch: bounded
|
|
29
29
|
internal reports to its Run and original driver.
|
|
30
30
|
- Sol pairs `axstack-auditor-sol`/`axstack-research-code-sol`/
|
|
31
|
-
`axstack-explore-execution-sol`: `codex/
|
|
31
|
+
`axstack-explore-execution-sol`: `codex/sol` high in
|
|
32
32
|
mixed/codex-only; intentionally absent in claude-only. Dispatch each
|
|
33
33
|
independently from its Sonnet seat on the same bounded brief without
|
|
34
34
|
cross-reading. The driver reconciles findings per claim, never averages.
|
|
@@ -11,22 +11,33 @@ skills root, or an explicit user selection in the run record. Missing or contrad
|
|
|
11
11
|
a setup gap: hold. Never infer from live profiles or `list_profiles`, harness,
|
|
12
12
|
tools, credentials, quota, subscription, or default to `mixed`.
|
|
13
13
|
|
|
14
|
-
At start, snapshot all 32 role IDs with provider/model/mode/effort; absent
|
|
14
|
+
At start, snapshot all 32 role IDs with provider/modelClass/model/mode/effort; absent
|
|
15
15
|
or unconfigured roles are recorded explicitly; never default.
|
|
16
16
|
Such a role holds only its work. Later installed or changed roles need an
|
|
17
17
|
explicit user decision to enter the snapshot. Live profiles
|
|
18
18
|
are authoritative at snapshot time and for availability; bundled presets are setup
|
|
19
19
|
inputs, not runtime proof.
|
|
20
|
+
For each role record class, resolved exact ID, source (catalog, transcript, or
|
|
21
|
+
pin), and time. Codex classes resolve through
|
|
22
|
+
`skills/axstack/scripts/resolve-models.js` with an explicit catalog
|
|
23
|
+
path; missing or malformed catalog holds. Claude classes start as `alias,
|
|
24
|
+
unresolved` until transcript read-back. Resume must reuse the snapshot and
|
|
25
|
+
never re-resolve it.
|
|
20
26
|
|
|
21
27
|
Preset changes apply to new runs only; an active run keeps its snapshot.
|
|
22
28
|
Changing it or replacing a session needs an explicit user decision and
|
|
23
29
|
revalidation. Unavailable models, efforts, roles, or overrides hold only affected
|
|
24
30
|
work; no automatic fallback, quota routing, subscription inference, or silent
|
|
25
|
-
provider/model/effort substitution.
|
|
31
|
+
provider/model/effort substitution. Only
|
|
32
|
+
explicit model rejection before the first turn permits Codex `--retry-of` with
|
|
33
|
+
the next eligible ID in the same class, provider, and effort. Fence the failed
|
|
34
|
+
Dispatch and record tried ID, error, and fallback ID in the snapshot and reply.
|
|
35
|
+
Timeout, quota, auth, and other failures hold; Claude rejection holds.
|
|
26
36
|
|
|
27
37
|
Load the [Role roster](role-roster.md) for configured roles and authored-review pairings.
|
|
28
38
|
|
|
29
|
-
Provenance is matched on provider/model
|
|
39
|
+
Provenance is matched on provider/model class derived from the recorded exact ID;
|
|
40
|
+
effort never maps. Missing table-row
|
|
30
41
|
provenance is unsupported and `INCOMPLETE`; report it and ask the user. Never
|
|
31
42
|
infer from slot, driver, owner, or provider. Author and owner never review.
|
|
32
43
|
|
|
@@ -75,7 +86,7 @@ reason in the run record, or in the brief for tiny direct work.
|
|
|
75
86
|
Require an approved spec plus a ticket map tied to that exact spec
|
|
76
87
|
revision, with acceptance checks and dependencies in the explicitly selected
|
|
77
88
|
Markdown, GitHub Issues, or Linear store. Prepare via `axstack-align` -> `axstack-spec`
|
|
78
|
-
(one approval) -> `axstack-tickets` -> handoff, then
|
|
89
|
+
(one approval) -> `axstack-tickets` -> handoff, then continue under autopilot when eligible.
|
|
79
90
|
- **Small:** clear, bounded one-PR work. The driver captures the named
|
|
80
91
|
**small-change intent** from the current request or user-chosen existing
|
|
81
92
|
issue plus explicit acceptance checks and exclusions, snapshots it once, and
|
|
@@ -96,7 +107,7 @@ not alone a formal spec trigger. Hold affected unsafe work while reassessing.
|
|
|
96
107
|
## Lifecycle routes (mode-specific scope identity required)
|
|
97
108
|
|
|
98
109
|
- Preparation: substantial work follows the align -> spec -> tickets ->
|
|
99
|
-
handoff path above, then
|
|
110
|
+
handoff path above, then continues under autopilot when eligible; small work uses the driver-captured
|
|
100
111
|
small-change intent.
|
|
101
112
|
- Execution: with its identity present, `axstack-implement` ->
|
|
102
113
|
`axstack-review` -> `axstack-watch`.
|
|
@@ -151,6 +151,8 @@ Authority: <who authorized which mutation>
|
|
|
151
151
|
Intent: <approved spec rev | small-change intent | adopted snapshot | peer/read-only mode>
|
|
152
152
|
Routing: <preset + source + snapshot ref>
|
|
153
153
|
Notification policy: <none | transport/target label/host/instructions path>
|
|
154
|
+
Autopilot: on | paused (<hold>; resume: <condition>) | off (cancelled <ts>); next: <step>
|
|
155
|
+
Release: <AGENTS.md file:line + tag-triggered workflow path + named install hosts> | not applicable (<reason>)
|
|
154
156
|
Source base: <exact revision or source identity>
|
|
155
157
|
IDs: <repo/project + workspace/agent receipt pointers>
|
|
156
158
|
Worktrees in other repositories: <per-run repository and worktree IDs or none>
|
|
@@ -0,0 +1,46 @@
|
|
|
1
|
+
import { readFileSync } from 'node:fs';
|
|
2
|
+
|
|
3
|
+
const args = process.argv.slice(2);
|
|
4
|
+
const option = (name) => {
|
|
5
|
+
const index = args.indexOf(name);
|
|
6
|
+
return index < 0 ? null : args[index + 1];
|
|
7
|
+
};
|
|
8
|
+
const path = option('--catalog');
|
|
9
|
+
const modelClass = option('--class');
|
|
10
|
+
const effort = option('--effort');
|
|
11
|
+
|
|
12
|
+
try {
|
|
13
|
+
if (!path || !['astra', 'sol', 'luna'].includes(modelClass) || !effort) {
|
|
14
|
+
throw new Error('expected --catalog path --class astra|sol|luna --effort level');
|
|
15
|
+
}
|
|
16
|
+
const catalog = JSON.parse(readFileSync(path, 'utf8'));
|
|
17
|
+
if (!Array.isArray(catalog.models) || typeof catalog.client_version !== 'string'
|
|
18
|
+
|| typeof catalog.fetched_at !== 'string') {
|
|
19
|
+
throw new Error('malformed catalog');
|
|
20
|
+
}
|
|
21
|
+
const classPattern = new RegExp(`^gpt-(\\d+(?:\\.\\d+)*)-${modelClass}$`);
|
|
22
|
+
const candidates = catalog.models
|
|
23
|
+
.filter((entry) => entry && entry.visibility === 'list'
|
|
24
|
+
&& typeof entry.slug === 'string'
|
|
25
|
+
&& classPattern.test(entry.slug)
|
|
26
|
+
&& Array.isArray(entry.supported_reasoning_levels)
|
|
27
|
+
&& entry.supported_reasoning_levels.some((level) => level?.effort === effort))
|
|
28
|
+
.map((entry) => entry.slug)
|
|
29
|
+
.sort((left, right) => {
|
|
30
|
+
const a = left.slice(4, -(modelClass.length + 1)).split('.').map(Number);
|
|
31
|
+
const b = right.slice(4, -(modelClass.length + 1)).split('.').map(Number);
|
|
32
|
+
for (let i = 0; i < Math.max(a.length, b.length); i++) {
|
|
33
|
+
const difference = (b[i] ?? 0) - (a[i] ?? 0);
|
|
34
|
+
if (difference) return difference;
|
|
35
|
+
}
|
|
36
|
+
return 0;
|
|
37
|
+
});
|
|
38
|
+
if (!candidates.length) throw new Error(`no eligible ${modelClass} model in catalog`);
|
|
39
|
+
console.log(JSON.stringify({
|
|
40
|
+
provider: 'codex', modelClass, effort, model: candidates[0], candidates,
|
|
41
|
+
catalog: { path, client_version: catalog.client_version, fetched_at: catalog.fetched_at },
|
|
42
|
+
}));
|
|
43
|
+
} catch (error) {
|
|
44
|
+
console.error(`model catalog resolution hold: ${error.message}`);
|
|
45
|
+
process.exitCode = 1;
|
|
46
|
+
}
|
|
@@ -5,6 +5,9 @@ description: When exploring or planning engineering work, use axstack-align to s
|
|
|
5
5
|
|
|
6
6
|
# Align
|
|
7
7
|
|
|
8
|
+
For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
|
|
9
|
+
for phase continuation and holds.
|
|
10
|
+
|
|
8
11
|
On driver entry, sweep under [Workspace hygiene](../axstack/references/workspace-hygiene.md); dispatched workers do not sweep.
|
|
9
12
|
For every dispatch brief, name its private `<run dir>/evidence/<dispatch>/` folder.
|
|
10
13
|
|
|
@@ -166,7 +169,7 @@ record or spec. Read-only scope keeps proposed documentation in the permitted
|
|
|
166
169
|
private record or response. Documentation is neither implementation nor spec
|
|
167
170
|
approval; record chosen document names and paths once per run.
|
|
168
171
|
|
|
169
|
-
## Read back, classify, and
|
|
172
|
+
## Read back, classify, and route
|
|
170
173
|
|
|
171
174
|
1. Read back the decisions, constraints, exclusions, and remaining evidence
|
|
172
175
|
gaps. For substantial work, this summary becomes part of the draft spec in
|
|
@@ -189,6 +192,7 @@ approval; record chosen document names and paths once per run.
|
|
|
189
192
|
[Orca runtime](../axstack/references/orca-runtime.md) immediately before
|
|
190
193
|
actual dispatch. Alignment completion never dispatches a recipient.
|
|
191
194
|
|
|
192
|
-
Alignment
|
|
193
|
-
identity is explicit
|
|
194
|
-
|
|
195
|
+
Alignment completes for both sizes only when the handoff is usable and its next
|
|
196
|
+
scope identity is explicit. An eligible delivery run continues under Autopilot;
|
|
197
|
+
an explicit stop-after-Align request ends here. Substantial work continues to
|
|
198
|
+
Spec, and small work continues from its small-change intent to Implement.
|
|
@@ -28,8 +28,8 @@ immediately before an actual auditor profile or session dispatch. Ordinary
|
|
|
28
28
|
audit reading and record writing do not load it, and the auditor never
|
|
29
29
|
dispatches.
|
|
30
30
|
|
|
31
|
-
Core owns the `axstack-auditor` profile (claude/
|
|
32
|
-
mixed/claude-only; codex/
|
|
31
|
+
Core owns the `axstack-auditor` profile (claude/sonnet high in
|
|
32
|
+
mixed/claude-only; codex/luna xhigh in codex-only) and its invocation.
|
|
33
33
|
This skill governs what that auditor reads, measures, and proposes.
|
|
34
34
|
Dispatch `axstack-auditor` and `axstack-auditor-sol` independently on the same
|
|
35
35
|
bounded brief, without cross-reading. The driver reconciles findings per claim;
|
|
@@ -91,6 +91,14 @@ counts with denominators plus the evidence behind the count:
|
|
|
91
91
|
evidence path is absent, or `UNKNOWN` with the reason when its records are
|
|
92
92
|
unavailable.
|
|
93
93
|
- Independent exact-revision review status and unresolved findings.
|
|
94
|
+
For authored PRs, measure the selected reviewer against this class pairing:
|
|
95
|
+
|
|
96
|
+
| Preset | Author class | Reviewer (class/effort) |
|
|
97
|
+
| --- | --- | --- |
|
|
98
|
+
| `mixed` | `codex/sol` | `axstack-reviewer-secondary` (`claude/opus` medium) |
|
|
99
|
+
| `mixed` | `claude/opus` | `axstack-reviewer-primary` (`codex/sol` high) |
|
|
100
|
+
| `codex-only` | `codex/sol` | `axstack-reviewer-secondary` (`codex/luna` xhigh) |
|
|
101
|
+
| `claude-only` | `claude/opus` | `axstack-reviewer-secondary` (`claude/sonnet` high) |
|
|
94
102
|
- Simplification applicability determinations evidenced / total candidate diffs,
|
|
95
103
|
and complete simplification receipts / total candidates, broken down as
|
|
96
104
|
`applied`, `not-applicable`, or `UNKNOWN` with the reason. This measures
|
|
@@ -5,6 +5,9 @@ description: When an approved task is ready to build or repair, use axstack-impl
|
|
|
5
5
|
|
|
6
6
|
# Implement
|
|
7
7
|
|
|
8
|
+
For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
|
|
9
|
+
for phase continuation and holds.
|
|
10
|
+
|
|
8
11
|
On driver entry, sweep under [Workspace hygiene](../axstack/references/workspace-hygiene.md); dispatched workers do not sweep.
|
|
9
12
|
For every dispatch brief, name its private `<run dir>/evidence/<dispatch>/` folder.
|
|
10
13
|
Include [Safe deletion](../axstack/references/workspace-hygiene.md#safe-deletion) in author briefs.
|
|
@@ -215,23 +218,34 @@ For each PR:
|
|
|
215
218
|
`repairs` once and counts once toward the third-round hold.
|
|
216
219
|
|
|
217
220
|
One run-level completion wait covers every unsettled Dispatch; the bounded
|
|
218
|
-
forge check wait is the only other wait.
|
|
221
|
+
forge check wait is the only other implementation wait. The eligible run arms
|
|
222
|
+
one maintain-mode chat-run watch at its first published PR; that watch owns its
|
|
223
|
+
10-minute harness wake or Orca fallback. End a turn only when every required PR
|
|
219
224
|
is `merge-ready` or `held`. Under the recorded Notification policy,
|
|
220
|
-
`axstack-relay` sends only a serious risk immediately
|
|
221
|
-
operation
|
|
222
|
-
|
|
225
|
+
`axstack-relay` sends only a serious risk immediately, a genuine blocked
|
|
226
|
+
operation needing user intervention after bounded safe recovery, or the
|
|
227
|
+
decision holds and capped milestones named by the recorded Notification policy.
|
|
228
|
+
Routine questions stay in Orca. Progress, CI pending, and completion always stay
|
|
223
229
|
in Orca.
|
|
230
|
+
Only the bounded categories—user-decision holds (including spec approval),
|
|
231
|
+
serious-risk holds, and at most two merge-ready/merged milestones per run—may
|
|
232
|
+
be relayed under the recorded Notification policy.
|
|
224
233
|
|
|
225
234
|
Merge-ready is the human boundary: the user merges, bottom-up for a stack. The
|
|
226
|
-
driver resumes on the user's next message
|
|
227
|
-
|
|
235
|
+
driver resumes on the user's next message, `/axstack-watch`, or the armed
|
|
236
|
+
chat-run watch wake; no Orca merge wake exists today.
|
|
237
|
+
Re-read forge state: record forge-merged PRs as `merged`;
|
|
228
238
|
changed heads or feedback return to step 1; retain useful author work before Close-out.
|
|
229
|
-
Run Close-out once only after every required PR is forge-merged
|
|
230
|
-
|
|
239
|
+
Run Close-out once only after every required PR is forge-merged, the run's
|
|
240
|
+
Release step is settled or not applicable, and acceptance passes. It settles
|
|
241
|
+
workers, records counts, makes the auditor decision and
|
|
231
242
|
settlement, releases worktrees, closes eligible tickets, and archives the run.
|
|
243
|
+
Without an Autopilot or Release record, the Release step is not applicable for
|
|
244
|
+
both Close-out and run completion.
|
|
232
245
|
|
|
233
246
|
The loop requires the `mixed` two-provider authored-review row. `codex-only` or
|
|
234
247
|
`claude-only` holds at step (3) for an explicit user routing choice, with no
|
|
235
248
|
substitution or same-provider review. Derived PR states are `authoring |
|
|
236
249
|
published | in-review | repairing(n) | merge-ready | merged | held`. The run is
|
|
237
|
-
done only when every required PR is forge-merged
|
|
250
|
+
done only when every required PR is forge-merged, the run's Release step is
|
|
251
|
+
settled or not applicable, and Close-out has receipts.
|
|
@@ -5,6 +5,9 @@ description: When the user requests a relay message or test, or an authorized no
|
|
|
5
5
|
|
|
6
6
|
# Relay
|
|
7
7
|
|
|
8
|
+
For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
|
|
9
|
+
for phase continuation and holds.
|
|
10
|
+
|
|
8
11
|
Send normal messages, transport tests, and authorized notifications to the
|
|
9
12
|
user through Hermes' native one-way `hermes send`. This is an inline caller
|
|
10
13
|
procedure: it creates no driver, team, owner, auditor, monitor, child session,
|
|
@@ -25,8 +28,12 @@ Choose the applicable message type:
|
|
|
25
28
|
in the caller's private notification policy. State the issue, impact, and the
|
|
26
29
|
answer or action needed.
|
|
27
30
|
- **Routine run events:** questions, spec approvals, progress, CI pending,
|
|
28
|
-
merge-ready, merged, and completion stay in Orca
|
|
29
|
-
|
|
31
|
+
merge-ready, merged, and completion stay in Orca unless the recorded
|
|
32
|
+
Notification policy names it. A policy may name only user-decision holds and
|
|
33
|
+
at most two merge-ready/merged milestones per run; deduplicate across implementation
|
|
34
|
+
and release. Progress, CI pending, and completion are never eligible merely
|
|
35
|
+
because a policy exists. They never become proactive relay messages merely
|
|
36
|
+
because the run is waiting.
|
|
30
37
|
|
|
31
38
|
Verify the transport, execution host, and intended recipient from the user's
|
|
32
39
|
request, trusted caller context, or an existing private notification policy.
|
|
@@ -204,17 +204,18 @@ This section applies to peer and authored PR modes.
|
|
|
204
204
|
- **Authored:** exactly one eligible independent reviewer from this complete
|
|
205
205
|
mapping:
|
|
206
206
|
|
|
207
|
-
| Preset | Actual author provider/
|
|
207
|
+
| Preset | Actual author provider/class | Reviewer role (configured class/effort) |
|
|
208
208
|
| --- | --- | --- |
|
|
209
|
-
| `mixed` |
|
|
210
|
-
| `mixed` |
|
|
211
|
-
| `codex-only` |
|
|
212
|
-
| `claude-only` |
|
|
209
|
+
| `mixed` | `codex/sol` | `axstack-reviewer-secondary` (`claude/opus` medium) |
|
|
210
|
+
| `mixed` | `claude/opus` | `axstack-reviewer-primary` (`codex/sol` high) |
|
|
211
|
+
| `codex-only` | `codex/sol` | `axstack-reviewer-secondary` (`codex/luna` xhigh) |
|
|
212
|
+
| `claude-only` | `claude/opus` | `axstack-reviewer-secondary` (`claude/sonnet` high) |
|
|
213
213
|
|
|
214
214
|
The diligence receipt is separate and does not count as a reviewer receipt.
|
|
215
215
|
|
|
216
|
-
|
|
217
|
-
effort to create a mapping.
|
|
216
|
+
From the recorded exact model ID, derive its class and match provenance
|
|
217
|
+
on provider/class; record effort, but never use effort to create a mapping.
|
|
218
|
+
An ID with no class is `INCOMPLETE`. Any other author provenance for the
|
|
218
219
|
selected preset is unsupported and `INCOMPLETE`, including its secondary
|
|
219
220
|
reviewer model, Astra, Luna, or Fable. Report the exact provenance gap and
|
|
220
221
|
ask the user. Never derive a reverse pairing from slot position. The
|
|
@@ -233,6 +234,11 @@ This section applies to peer and authored PR modes.
|
|
|
233
234
|
effort and spawn no redundant final reviewer. If a required reviewer is
|
|
234
235
|
unavailable, report that exact model gap, mark review `INCOMPLETE`, and ask
|
|
235
236
|
the user; do not lower effort or choose any automatic fallback.
|
|
237
|
+
The only within-class exception is explicit model rejection before the first
|
|
238
|
+
turn: Codex may use
|
|
239
|
+
`--retry-of` with the next eligible ID in the same class, provider, and
|
|
240
|
+
effort; fence and record the rejected attempt. Timeout, quota, and auth
|
|
241
|
+
failures hold; Claude rejection holds. Never cross class or provider.
|
|
236
242
|
|
|
237
243
|
Continue only when session receipts prove the required models, non-author
|
|
238
244
|
independence, actual author provenance where applicable, and exact brief.
|
|
@@ -5,6 +5,9 @@ description: When agreed work needs an approved baseline, use axstack-spec to wr
|
|
|
5
5
|
|
|
6
6
|
# Specification baseline
|
|
7
7
|
|
|
8
|
+
For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
|
|
9
|
+
for phase continuation and holds.
|
|
10
|
+
|
|
8
11
|
On driver entry, sweep under [Workspace hygiene](../axstack/references/workspace-hygiene.md); dispatched workers do not sweep.
|
|
9
12
|
For every dispatch brief, name its private `<run dir>/evidence/<dispatch>/` folder.
|
|
10
13
|
|
|
@@ -75,7 +78,8 @@ Material change: <none | description + affected PRs/tasks + hold state>
|
|
|
75
78
|
The snapshot is ready for ticketing when its authoritative revision,
|
|
76
79
|
counterpart, and preserved ref resolve to the approved content. Return that
|
|
77
80
|
exact identity; routine execution of the settled plan needs no repeat adviser
|
|
78
|
-
consultation or spec approval.
|
|
81
|
+
consultation or spec approval. In an eligible delivery run with no hold,
|
|
82
|
+
continue to Tickets in the same driver chat.
|
|
79
83
|
|
|
80
84
|
## Material revisions
|
|
81
85
|
|
|
@@ -5,12 +5,15 @@ description: When an approved capability needs executable tasks, use axstack-tic
|
|
|
5
5
|
|
|
6
6
|
# Tickets
|
|
7
7
|
|
|
8
|
+
For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
|
|
9
|
+
for phase continuation and holds.
|
|
10
|
+
|
|
8
11
|
On driver entry, sweep under [Workspace hygiene](../axstack/references/workspace-hygiene.md); dispatched workers do not sweep.
|
|
9
12
|
For every dispatch brief, name its private `<run dir>/evidence/<dispatch>/` folder.
|
|
10
13
|
|
|
11
14
|
Produce an executable capability map tied to the exact approved spec revision.
|
|
12
15
|
Keep user-visible capabilities in the selected store, keep implementation detail
|
|
13
|
-
in the repository, reconcile lifecycle state, and
|
|
16
|
+
in the repository, reconcile lifecycle state, and return a map for continuation.
|
|
14
17
|
|
|
15
18
|
Before mapping, load [Standing contracts](../axstack/references/contracts.md).
|
|
16
19
|
Follow its required edge to [Shared lifecycle](../axstack/references/lifecycle.md),
|
|
@@ -100,5 +103,5 @@ Recommendation: <move to In Review | keep open | close | other> (driver verifies
|
|
|
100
103
|
|
|
101
104
|
5. **Return the mapping.** Report the pinned spec revision, selected store, map
|
|
102
105
|
references, mutations performed by the driver, recorded gaps, and unresolved
|
|
103
|
-
decisions.
|
|
104
|
-
|
|
106
|
+
decisions. With a complete map and no hold, an eligible delivery run
|
|
107
|
+
continues to Implement in the same driver chat.
|
|
@@ -5,6 +5,9 @@ description: When babysitting an existing PR, use axstack-watch to monitor or ma
|
|
|
5
5
|
|
|
6
6
|
# Watch
|
|
7
7
|
|
|
8
|
+
For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
|
|
9
|
+
for phase continuation and holds.
|
|
10
|
+
|
|
8
11
|
On driver entry, sweep under [Workspace hygiene](../axstack/references/workspace-hygiene.md); dispatched workers do not sweep.
|
|
9
12
|
For every dispatch brief, name its private `<run dir>/evidence/<dispatch>/` folder.
|
|
10
13
|
|
|
@@ -49,7 +52,8 @@ authority is unverified, record the hold and continue read-only.
|
|
|
49
52
|
|
|
50
53
|
Choose one mode from the user's authority and record it before dispatch:
|
|
51
54
|
|
|
52
|
-
- **Chat-run watch:**
|
|
55
|
+
- **Chat-run watch:** authorized maintain mode is the default for run-created
|
|
56
|
+
PRs. The initiating chat remains the only driver and record
|
|
53
57
|
writer for every PR raised in its Run, including later verified publications
|
|
54
58
|
and explicitly adopted members. Follow [Chat-run watch runtime](references/watch-runtime.md#chat-run-watch)
|
|
55
59
|
for its scheduled driver wake and Orca fallback. This mode has no replacement `axstack-owner` or
|
|
@@ -121,11 +125,16 @@ A handled wake has an acknowledged event ID, an observation or action bound to
|
|
|
121
125
|
the current revision, and a recorded hold or next owner where work remains.
|
|
122
126
|
|
|
123
127
|
Under a recorded `Notification policy`, the owner may use the optional
|
|
124
|
-
[axstack-relay](../axstack-relay/SKILL.md) only for a serious risk immediately
|
|
125
|
-
|
|
126
|
-
recovery
|
|
127
|
-
|
|
128
|
-
|
|
128
|
+
[axstack-relay](../axstack-relay/SKILL.md) only for a serious risk immediately,
|
|
129
|
+
a genuine blocked operation needing user intervention after bounded safe
|
|
130
|
+
recovery, or decision holds and capped milestones named by the recorded policy.
|
|
131
|
+
Routine questions stay in Orca. Progress, CI pending, and completion always stay
|
|
132
|
+
in Orca.
|
|
133
|
+
Only the bounded categories—user-decision holds (including spec approval),
|
|
134
|
+
serious-risk holds, and at most two merge-ready/merged milestones per run—may
|
|
135
|
+
be relayed under the recorded Notification policy.
|
|
136
|
+
The standalone monitor never sends; the chat-run observer reports only
|
|
137
|
+
internally. Deduplicate authorized notifications;
|
|
129
138
|
absent policy or failed relay uses the current Orca conversation and leaves
|
|
130
139
|
the existing hold open.
|
|
131
140
|
|
|
@@ -148,10 +157,14 @@ human approval remain allowed.
|
|
|
148
157
|
|
|
149
158
|
## 6. End and preserve continuity
|
|
150
159
|
|
|
151
|
-
End a chat-run watch after all members merged or closed
|
|
152
|
-
|
|
153
|
-
|
|
154
|
-
|
|
160
|
+
End a chat-run watch after all members merged or closed and the run's release
|
|
161
|
+
step is settled or not applicable, user cancellation, or the recorded wake
|
|
162
|
+
expires. Without an Autopilot or Release record, the release step is not
|
|
163
|
+
applicable to this watch. A required PR closed without merging records a
|
|
164
|
+
decision hold and the wake remains active while unexpired until the user
|
|
165
|
+
resolves scope, cancels, or the wake expires. Stop the chosen wake and verify
|
|
166
|
+
its stop receipt; a failed or uncertain harness wake stop is a hold.
|
|
167
|
+
The Orca fallback also needs own-automation disable/readback and driver-owned automation
|
|
155
168
|
removal and workspace cleanup under
|
|
156
169
|
[Watch runtime](references/watch-runtime.md#chat-run-watch).
|
|
157
170
|
|
|
@@ -183,6 +196,6 @@ Resume: <known commands or verified refs needed to reconcile from this revision>
|
|
|
183
196
|
|
|
184
197
|
The watch ends only when registrations are stopped, receipts are recorded, and
|
|
185
198
|
the PR is either merged or represented by this resumable state.
|
|
186
|
-
When every required PR is merged
|
|
187
|
-
[Close-out](../axstack/references/lifecycle.md#close-out)
|
|
188
|
-
run as done.
|
|
199
|
+
When every required PR is merged and the run's Release step is settled or not
|
|
200
|
+
applicable, follow the lifecycle [Close-out](../axstack/references/lifecycle.md#close-out)
|
|
201
|
+
before reporting the run as done.
|