axstack 0.20.29 → 0.20.31
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +6 -4
- package/bin/axstack.js +1 -0
- package/docs/installation.md +20 -9
- package/docs/workflows.md +42 -17
- package/package.json +1 -1
- package/profiles/presets/claude-only.json +81 -45
- package/profiles/presets/codex-only.json +78 -42
- package/profiles/presets/mixed.json +83 -47
- package/skills/axstack/references/autopilot.md +108 -0
- package/skills/axstack/references/candidate-publication.md +9 -0
- package/skills/axstack/references/contracts.md +5 -1
- package/skills/axstack/references/diligence.md +23 -0
- package/skills/axstack/references/lifecycle.md +2 -2
- package/skills/axstack/references/orca-runtime.md +25 -6
- package/skills/axstack/references/role-roster.md +37 -0
- package/skills/axstack/references/routing.md +19 -31
- package/skills/axstack/references/run-record.md +2 -0
- package/skills/axstack/scripts/resolve-models.js +46 -0
- package/skills/axstack-align/SKILL.md +8 -4
- package/skills/axstack-audit/SKILL.md +15 -2
- package/skills/axstack-implement/SKILL.md +30 -11
- package/skills/axstack-improve/SKILL.md +5 -0
- package/skills/axstack-relay/SKILL.md +9 -2
- package/skills/axstack-research/SKILL.md +14 -2
- package/skills/axstack-review/SKILL.md +29 -13
- package/skills/axstack-spec/SKILL.md +8 -1
- package/skills/axstack-tickets/SKILL.md +9 -3
- package/skills/axstack-watch/SKILL.md +27 -13
- package/skills/axstack-watch/references/watch-runtime.md +11 -4
- package/src/installer.js +8 -0
- package/src/roles.js +39 -11
|
@@ -0,0 +1,108 @@
|
|
|
1
|
+
# Autopilot
|
|
2
|
+
|
|
3
|
+
This reference applies to authorized engineering-delivery runs. The original
|
|
4
|
+
driver remains the sole run-record writer and phase router in the same chat;
|
|
5
|
+
phase completion is not a native ownership handoff. Explicit planning-only,
|
|
6
|
+
read-only, stop-after-phase, observation-only, and peer requests retain their
|
|
7
|
+
selected boundary. A status question such as "what's left" is observation,
|
|
8
|
+
not a mode change.
|
|
9
|
+
|
|
10
|
+
## Advance and hold
|
|
11
|
+
|
|
12
|
+
Advance only after the finishing phase returns its completed identity (a
|
|
13
|
+
small-change intent, approved spec, matching ticket map, merge-ready or merged
|
|
14
|
+
state) and the run record has no open hold. A hold from any phase stops the run:
|
|
15
|
+
record its reason, owner, and resume condition, then take no dependent action.
|
|
16
|
+
That covers tracker access, adviser or arena-seat availability, diligence
|
|
17
|
+
FINDINGS when the phase records a hold, CI-wait timeout, readiness UNKNOWN,
|
|
18
|
+
dismissed approval, wake or cleanup uncertainty, single-provider routing, an
|
|
19
|
+
existing tag or version, and failed publish. Diligence FINDINGS during implement
|
|
20
|
+
follow its §6 repair route; at spec, tickets, or release preparation the driver
|
|
21
|
+
resolves them before advancing, and only a recorded hold pauses autopilot.
|
|
22
|
+
|
|
23
|
+
Record `Autopilot: on | paused (<hold>; resume: <condition>) | off (cancelled
|
|
24
|
+
<ts>)` and the next step in the private run record. A user answer to the hold
|
|
25
|
+
resumes after reconciliation; silence does not.
|
|
26
|
+
Awaiting human spec approval records `Autopilot: paused (spec approval; resume:
|
|
27
|
+
human approval)` as a decision hold eligible under the Notification policy.
|
|
28
|
+
|
|
29
|
+
## Phase sequence
|
|
30
|
+
|
|
31
|
+
- Small: Align read-back, small-change intent, implement, watch in maintain
|
|
32
|
+
mode, human merge. An opted-in Align refinement is part of read-back.
|
|
33
|
+
- Substantial: Align, spec draft with advisers and diligence, human spec
|
|
34
|
+
approval at gate 1, tickets with diligence, implement, watch in maintain mode,
|
|
35
|
+
merge-ready, human merge. An opted-in Align refinement is part of gate 1.
|
|
36
|
+
|
|
37
|
+
Do not seek another phase-start instruction after a completed identity.
|
|
38
|
+
Spec approval is always the human's decision. Every PR merge is the human's,
|
|
39
|
+
including a release PR and each PR in a stack, bottom-up.
|
|
40
|
+
|
|
41
|
+
## Implement into maintain watch
|
|
42
|
+
|
|
43
|
+
When implement publishes the run's first PR, arm exactly one `axstack-watch`
|
|
44
|
+
chat-run in authorized maintain mode. Use the 10-minute harness wake, with the
|
|
45
|
+
existing Orca fallback when unavailable. Later run PRs join after verified
|
|
46
|
+
publication readback; an explicitly adopted PR joins only with its maintenance
|
|
47
|
+
snapshot. The original driver alone routes work; one author writes each
|
|
48
|
+
candidate. Until a PR is merge-ready, wakes feed implement §6 step 4. After
|
|
49
|
+
merge-ready, watch §5 maintenance repairs feedback, rebases when the base moves,
|
|
50
|
+
keeps CI green, and checks approvals without re-requesting human review.
|
|
51
|
+
|
|
52
|
+
Maintain is the default mode for run-created PRs. End the chat-run watch when
|
|
53
|
+
every watched PR is merged or closed and the run's release step is settled or
|
|
54
|
+
not applicable, or when the user cancels. Expiry is a recorded stop with
|
|
55
|
+
resumable state, never a silent renewal. A required PR closed without merging
|
|
56
|
+
is incomplete scope; it does not make the run release-eligible. On wake expiry
|
|
57
|
+
record `Autopilot: paused (wake expired; resume: user reauthorizes a wake)` and
|
|
58
|
+
notify under the recorded Notification policy when user action is needed.
|
|
59
|
+
|
|
60
|
+
## Release and install, when applicable
|
|
61
|
+
|
|
62
|
+
Detect applicability once at Align or spec time. Record `Release: <AGENTS.md
|
|
63
|
+
file:line + tag-triggered workflow path + named install hosts> | not applicable
|
|
64
|
+
(<reason>)`. The predicate is an AGENTS.md release rule naming an existing
|
|
65
|
+
tag-triggered workflow. A partial match is not applicable and its reason is
|
|
66
|
+
noted. Install hosts come only from explicit targets; an absent host list is a
|
|
67
|
+
decision hold, not permission to infer hosts. A missing install host list at
|
|
68
|
+
Align or spec time is a decision hold before release authority is presented.
|
|
69
|
+
|
|
70
|
+
Show the `Release:` line in the spec for human approval at gate 1, or the small
|
|
71
|
+
work Align read-back. Copy that decision to `Authority:` in the run record.
|
|
72
|
+
This authority is per run and never carries over to another run or repository.
|
|
73
|
+
The small-work Align read-back names the existing Release and host-mutation
|
|
74
|
+
authority and explicit hosts; silence cannot fill a missing authority or target.
|
|
75
|
+
|
|
76
|
+
After all required feature PRs merge, open one release PR. Default to a patch
|
|
77
|
+
version, or minor if a `feat` commit landed since the last tag. This normal run
|
|
78
|
+
PR gets authored review and diligence of its body against merged PRs, reaches
|
|
79
|
+
merge-ready, then waits for human merge. Once the forge confirms that merge,
|
|
80
|
+
tag and wait for the staged publish. Human npm stage approval is a decision
|
|
81
|
+
hold: agents never run `npm stage approve`. A wake verifies the registry reports
|
|
82
|
+
the expected package and version. Install on the named hosts, verify version
|
|
83
|
+
and roles, then run Close-out last with release and install receipts and the
|
|
84
|
+
installed version.
|
|
85
|
+
|
|
86
|
+
An existing version or tag, failed publish, pending approval, uncertain
|
|
87
|
+
registry result, missing host access, or failed install verification is a
|
|
88
|
+
resumable hold, never success. Tagging, publishing, installation, and host
|
|
89
|
+
mutation require the recorded per-run authority and their existing checks.
|
|
90
|
+
|
|
91
|
+
## Resume, cancel, and notify
|
|
92
|
+
|
|
93
|
+
At every entry (user message, wake, compaction, or new chat), reconcile the
|
|
94
|
+
owner, authoritative Dispatch, approved revision, PR membership, uncertain
|
|
95
|
+
tags, wakes, publications, and completed receipts under lifecycle and
|
|
96
|
+
run-record before advancing. Only the original driver advances. Wakes do not
|
|
97
|
+
reset attempt budgets and do not grant approvals. Cancel sets `Autopilot: off`,
|
|
98
|
+
stops new actions, and ends the watch under watch §6 with guarded settlement.
|
|
99
|
+
Cancellation does not cancel a running author Dispatch by inference; let it
|
|
100
|
+
report, then settle that exact Dispatch under lifecycle guards without new
|
|
101
|
+
publication.
|
|
102
|
+
|
|
103
|
+
Use the run's recorded Notification policy through `axstack-relay`.
|
|
104
|
+
Decision holds, including spec and npm approval, are always eligible. Across
|
|
105
|
+
implementation and release, merge-ready and merged notifications together are
|
|
106
|
+
capped at two per run; deduplicate by purpose and revision. Healthy ticks stay
|
|
107
|
+
quiet. A failed or uncertain delivery preserves the underlying hold. A relay
|
|
108
|
+
message is only a notification, never authority to approve, merge, or publish.
|
|
@@ -6,6 +6,11 @@ Within recorded PR-scoped publication authority, the owner reconciles that
|
|
|
6
6
|
receipt against the actual local candidate SHA and base. The owner does not edit
|
|
7
7
|
the author's candidate; required code changes return to the author.
|
|
8
8
|
|
|
9
|
+
Before publication, dispatch `axstack-diligence` under
|
|
10
|
+
[Diligence](diligence.md) to check the author receipt against its evidence
|
|
11
|
+
folder: red/green logs exist, and counts, SHAs, and paths match. Resolve
|
|
12
|
+
`FINDINGS` with the same author before publishing.
|
|
13
|
+
|
|
9
14
|
Publish the existing commits through `gh stack`. Prefer a fast-forward push.
|
|
10
15
|
Before a history rewrite, confirm the expected-old remote SHA and use lease
|
|
11
16
|
protection; a mismatch holds publication. If the push outcome is ambiguous,
|
|
@@ -40,3 +45,7 @@ as a `git clone` into a temp directory followed by `orca repo add`; each
|
|
|
40
45
|
the directory is gone. Release preparation uses a `release/<version>` worktree
|
|
41
46
|
of the same registered repo the same way. Release the checkout with
|
|
42
47
|
`ORCA worktree rm` after its receipt is recorded.
|
|
48
|
+
|
|
49
|
+
For a release PR, dispatch `axstack-diligence` under
|
|
50
|
+
[Diligence](diligence.md) to check the release PR body
|
|
51
|
+
against the merged PRs before publication.
|
|
@@ -40,7 +40,11 @@ Validate the configured provider and model at actual launch. If it is
|
|
|
40
40
|
unavailable or exhausted, pause affected work, record the gap, and ask the
|
|
41
41
|
user. Never infer a route from quota state or subscription entitlement. Every
|
|
42
42
|
substitution requires the user's decision: configured alternatives and native
|
|
43
|
-
fallback prose are not defaults.
|
|
43
|
+
fallback prose are not defaults. The only within-class exception is explicit
|
|
44
|
+
model rejection before the first turn: Codex may retry with `--retry-of` using
|
|
45
|
+
the next eligible version in the same class, provider, and effort, recording
|
|
46
|
+
the failed ID, error, and fallback ID. Claude rejection holds. Timeout, quota,
|
|
47
|
+
and auth failures hold.
|
|
44
48
|
|
|
45
49
|
## Driver and adviser split
|
|
46
50
|
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
# Diligence
|
|
2
|
+
|
|
3
|
+
Dispatch `axstack-diligence` through Orca with a pinned brief and evidence paths.
|
|
4
|
+
It is read-only, never authors or edits, and returns `PASS` or `FINDINGS`
|
|
5
|
+
with locations, observed evidence, and limits. A stale or missing receipt is
|
|
6
|
+
not a pass. Keep its first pass independent of other reviewers and workers.
|
|
7
|
+
|
|
8
|
+
For a PR, compare every changed line with the accepted intent and exclusions:
|
|
9
|
+
is it intended and in scope? Check that no contract, rule, or obligation was
|
|
10
|
+
silently weakened or dropped by rewording. Compare the PR body, commit messages,
|
|
11
|
+
and author receipt with the diff: numbers, IDs, versions, test counts, sizes,
|
|
12
|
+
paths, and stale references. Bind the result to the exact head and base.
|
|
13
|
+
|
|
14
|
+
For research, reopen cited sources for answer-changing claims before the
|
|
15
|
+
driver folds verified claims. For a draft spec, compare it with Align decisions
|
|
16
|
+
before user approval: flag anything dropped, added, or softened. For tickets,
|
|
17
|
+
map every spec acceptance item to a capability's acceptance. Before candidate
|
|
18
|
+
publication, compare the author receipt with its evidence folder: red/green
|
|
19
|
+
logs exist, and counts, SHAs, and paths match. For release preparation, compare
|
|
20
|
+
the release PR body with the merged PRs.
|
|
21
|
+
|
|
22
|
+
`FINDINGS` identifies a mismatch for the driver to resolve at the owning phase;
|
|
23
|
+
it does not edit the artifact or create another review round by itself.
|
|
@@ -107,8 +107,8 @@ Tracking grants no merge, release, model-substitution, or scope authority.
|
|
|
107
107
|
|
|
108
108
|
The default 24-hour deadline covers standalone task-owned timers. Stop them at
|
|
109
109
|
deadline and preserve remaining work; the review automation has no task-owned
|
|
110
|
-
deadline.
|
|
111
|
-
|
|
110
|
+
deadline. Merge-ready requires applicable review receipt(s) and current diligence
|
|
111
|
+
`PASS` at the exact head; CI/tests alone are insufficient. Merge-ready
|
|
112
112
|
differs from merged; human merges.
|
|
113
113
|
|
|
114
114
|
## Review automation health
|
|
@@ -38,23 +38,42 @@ Read `roles.json` from the installed shared root `skills/axstack/`. The installe
|
|
|
38
38
|
shape is `{ "version": 1, "preset": "<name>", "roles": [...] }`. Bundled
|
|
39
39
|
profiles are setup inputs shaped as
|
|
40
40
|
`{ "version": 1, "roles": [...] }`. A new run records the selected preset and
|
|
41
|
-
all
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
41
|
+
all 32 role rows once. For each role record class, resolved exact ID, source,
|
|
42
|
+
and time. An active run keeps the exact snapshot; resume reuses it without
|
|
43
|
+
re-resolution until the user explicitly changes it.
|
|
44
|
+
|
|
45
|
+
Select the requested role by stable ID. A missing class and missing or null
|
|
46
|
+
model holds only that role; never launch a provider default. Resolve Codex
|
|
47
|
+
classes with `scripts/resolve-models.js`, passing the catalog path explicitly;
|
|
48
|
+
missing or malformed catalogs hold. The first launch of each Claude class uses
|
|
49
|
+
its alias. Read the exact ID from the first assistant turn's `message.model` in
|
|
50
|
+
that worker's own session transcript at
|
|
51
|
+
`~/.claude/projects/<worktree-path-slug>/*.jsonl`; the worktree path slug
|
|
52
|
+
replaces each non-alphanumeric character with `-`. Identify the file by the
|
|
53
|
+
worker's session ID, or use the newest file created after launch. Later launches
|
|
54
|
+
of that class use the recorded exact ID. Before read-back record `alias,
|
|
55
|
+
unresolved`; record an unknown read-back as unknown and hold
|
|
56
|
+
provenance-dependent work. A worker self-report is a labeled last
|
|
57
|
+
resort. Launch-by-agent-id routes for which Orca exposes no
|
|
46
58
|
`--model` override (today: `grok`, `antigravity`) record `model: null` with an explicit note and are
|
|
47
59
|
launchable; the run record snapshots the model the TUI reports. Validate provider, model, and effort
|
|
48
60
|
against the guide and actual launch capability. Stored `modeId` and other
|
|
49
61
|
permission fields are conservative intent, not proof of effective permission
|
|
50
62
|
parity or a security boundary. Requested settings, input acceptance, effective
|
|
51
63
|
settings, and completed work are separate evidence. An unsupported or
|
|
52
|
-
unavailable value holds affected work for the user's decision
|
|
64
|
+
unavailable value holds affected work for the user's decision except the narrow
|
|
65
|
+
retry below.
|
|
53
66
|
The single-provider preset's null adviser and round-2 seat are intentional installation data, not
|
|
54
67
|
readiness failure; because Align and Spec require both adviser receipts, either
|
|
55
68
|
null adviser still holds those phases. The current chat is the driver and has
|
|
56
69
|
no role row in any preset.
|
|
57
70
|
|
|
71
|
+
Only explicit model rejection before the first turn permits a Codex
|
|
72
|
+
`--retry-of` with the next eligible ID in the same class, provider, and effort.
|
|
73
|
+
Fence the rejected Dispatch and record tried ID, error, and fallback ID in the
|
|
74
|
+
snapshot and reply. Timeout, quota, auth, and other failures hold; Claude
|
|
75
|
+
rejection holds. Apply this to every role, including advisers and judges.
|
|
76
|
+
|
|
58
77
|
## Materialize checkouts as worktrees of the registered repo
|
|
59
78
|
|
|
60
79
|
Every reviewer, release, or worker checkout is `ORCA worktree create --repo
|
|
@@ -0,0 +1,37 @@
|
|
|
1
|
+
# Role roster
|
|
2
|
+
|
|
3
|
+
- Chat drives (no role ID); `axstack-owner` owns one PR and
|
|
4
|
+
`axstack-author` its sole writer.
|
|
5
|
+
- `axstack-reviewer-primary` and `axstack-reviewer-secondary` are the ordered
|
|
6
|
+
peer pair. Peer review uses both; authored review uses this table:
|
|
7
|
+
|
|
8
|
+
| Preset | Author class | Reviewer (class/effort) |
|
|
9
|
+
| --- | --- | --- |
|
|
10
|
+
| `mixed` | `codex/sol` | `axstack-reviewer-secondary` (`claude/opus` medium) |
|
|
11
|
+
| `mixed` | `claude/opus` | `axstack-reviewer-primary` (`codex/sol` high) |
|
|
12
|
+
| `codex-only` | `codex/sol` | `axstack-reviewer-secondary` (`codex/luna` xhigh) |
|
|
13
|
+
| `claude-only` | `claude/opus` | `axstack-reviewer-secondary` (`claude/sonnet` high) |
|
|
14
|
+
- `axstack-advisor-astra`/`axstack-advisor-opus` advise and author candidates;
|
|
15
|
+
`axstack-arena-candidate-grok`/
|
|
16
|
+
`axstack-arena-candidate-antigravity` add families.
|
|
17
|
+
`axstack-arena-judge-opus` judges round 1; `axstack-escalation-fable`/`axstack-arena-judge-astra` judge round 2.
|
|
18
|
+
High-stakes/trigger: fresh [contract](contracts.md) session.
|
|
19
|
+
`axstack-auditor` audits; `axstack-checker` reports discrepancies.
|
|
20
|
+
- `axstack-explainer`/`axstack-explainer-review`: explain/review.
|
|
21
|
+
- `axstack-diligence`: read-only [diligence checks](diligence.md) for every PR
|
|
22
|
+
review round and bounded research, spec, ticket, receipt, and release claims.
|
|
23
|
+
- `axstack-ui-verifier`: [UI checks](ui-verification.md).
|
|
24
|
+
- `axstack-auditor`/`axstack-research-requirements`/
|
|
25
|
+
`axstack-research-code`/`axstack-research-web`/
|
|
26
|
+
`axstack-explore-execution`/`axstack-monitor`:
|
|
27
|
+
`claude/sonnet` high in mixed/claude-only.
|
|
28
|
+
`axstack-monitor`: standalone watch never sends; chat-run watch: bounded
|
|
29
|
+
internal reports to its Run and original driver.
|
|
30
|
+
- Sol pairs `axstack-auditor-sol`/`axstack-research-code-sol`/
|
|
31
|
+
`axstack-explore-execution-sol`: `codex/sol` high in
|
|
32
|
+
mixed/codex-only; intentionally absent in claude-only. Dispatch each
|
|
33
|
+
independently from its Sonnet seat on the same bounded brief without
|
|
34
|
+
cross-reading. The driver reconciles findings per claim, never averages.
|
|
35
|
+
Record intentional absence and continue with Sonnet alone; a configured
|
|
36
|
+
but unavailable seat holds only its affected work.
|
|
37
|
+
- `axstack-debug-investigator-1..4` probe L1 briefs.
|
|
@@ -11,45 +11,33 @@ skills root, or an explicit user selection in the run record. Missing or contrad
|
|
|
11
11
|
a setup gap: hold. Never infer from live profiles or `list_profiles`, harness,
|
|
12
12
|
tools, credentials, quota, subscription, or default to `mixed`.
|
|
13
13
|
|
|
14
|
-
At start, snapshot all
|
|
14
|
+
At start, snapshot all 32 role IDs with provider/modelClass/model/mode/effort; absent
|
|
15
15
|
or unconfigured roles are recorded explicitly; never default.
|
|
16
16
|
Such a role holds only its work. Later installed or changed roles need an
|
|
17
17
|
explicit user decision to enter the snapshot. Live profiles
|
|
18
18
|
are authoritative at snapshot time and for availability; bundled presets are setup
|
|
19
19
|
inputs, not runtime proof.
|
|
20
|
+
For each role record class, resolved exact ID, source (catalog, transcript, or
|
|
21
|
+
pin), and time. Codex classes resolve through
|
|
22
|
+
`skills/axstack/scripts/resolve-models.js` with an explicit catalog
|
|
23
|
+
path; missing or malformed catalog holds. Claude classes start as `alias,
|
|
24
|
+
unresolved` until transcript read-back. Resume must reuse the snapshot and
|
|
25
|
+
never re-resolve it.
|
|
20
26
|
|
|
21
27
|
Preset changes apply to new runs only; an active run keeps its snapshot.
|
|
22
28
|
Changing it or replacing a session needs an explicit user decision and
|
|
23
29
|
revalidation. Unavailable models, efforts, roles, or overrides hold only affected
|
|
24
30
|
work; no automatic fallback, quota routing, subscription inference, or silent
|
|
25
|
-
provider/model/effort substitution.
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
| `mixed` | Claude / Opus (`claude/claude-opus-5-5`) | `axstack-reviewer-primary` (`codex/gpt-6-sol` high) |
|
|
36
|
-
| `codex-only` | Codex / Sol (`codex/gpt-6-sol`) | `axstack-reviewer-secondary` (`codex/gpt-6-luna` xhigh) |
|
|
37
|
-
| `claude-only` | Claude / Opus (`claude/claude-opus-5-5`) | `axstack-reviewer-secondary` (`claude/claude-sonnet-5-5` high) |
|
|
38
|
-
- `axstack-advisor-astra`/`axstack-advisor-opus` advise and author candidates;
|
|
39
|
-
`axstack-arena-candidate-grok`/
|
|
40
|
-
`axstack-arena-candidate-antigravity` add families.
|
|
41
|
-
`axstack-arena-judge-opus` judges round 1; `axstack-escalation-fable`/`axstack-arena-judge-astra` judge round 2.
|
|
42
|
-
High-stakes/trigger: fresh [contract](contracts.md) session.
|
|
43
|
-
`axstack-auditor` audits; `axstack-checker` reports discrepancies.
|
|
44
|
-
- `axstack-explainer`/`axstack-explainer-review`: explain/review.
|
|
45
|
-
- `axstack-ui-verifier`: [UI checks](ui-verification.md).
|
|
46
|
-
- `axstack-research-requirements`/`axstack-research-web`/`axstack-monitor`:
|
|
47
|
-
Sonnet 5.5 high in mixed/claude-only.
|
|
48
|
-
`axstack-monitor`: standalone watch never sends; chat-run watch: bounded
|
|
49
|
-
internal reports to its Run and original driver.
|
|
50
|
-
- `axstack-debug-investigator-1..4` probe L1 briefs.
|
|
51
|
-
|
|
52
|
-
Provenance is matched on provider/model ID; effort never maps. Missing table-row
|
|
31
|
+
provider/model/effort substitution. Only
|
|
32
|
+
explicit model rejection before the first turn permits Codex `--retry-of` with
|
|
33
|
+
the next eligible ID in the same class, provider, and effort. Fence the failed
|
|
34
|
+
Dispatch and record tried ID, error, and fallback ID in the snapshot and reply.
|
|
35
|
+
Timeout, quota, auth, and other failures hold; Claude rejection holds.
|
|
36
|
+
|
|
37
|
+
Load the [Role roster](role-roster.md) for configured roles and authored-review pairings.
|
|
38
|
+
|
|
39
|
+
Provenance is matched on provider/model class derived from the recorded exact ID;
|
|
40
|
+
effort never maps. Missing table-row
|
|
53
41
|
provenance is unsupported and `INCOMPLETE`; report it and ask the user. Never
|
|
54
42
|
infer from slot, driver, owner, or provider. Author and owner never review.
|
|
55
43
|
|
|
@@ -98,7 +86,7 @@ reason in the run record, or in the brief for tiny direct work.
|
|
|
98
86
|
Require an approved spec plus a ticket map tied to that exact spec
|
|
99
87
|
revision, with acceptance checks and dependencies in the explicitly selected
|
|
100
88
|
Markdown, GitHub Issues, or Linear store. Prepare via `axstack-align` -> `axstack-spec`
|
|
101
|
-
(one approval) -> `axstack-tickets` -> handoff, then
|
|
89
|
+
(one approval) -> `axstack-tickets` -> handoff, then continue under autopilot when eligible.
|
|
102
90
|
- **Small:** clear, bounded one-PR work. The driver captures the named
|
|
103
91
|
**small-change intent** from the current request or user-chosen existing
|
|
104
92
|
issue plus explicit acceptance checks and exclusions, snapshots it once, and
|
|
@@ -119,7 +107,7 @@ not alone a formal spec trigger. Hold affected unsafe work while reassessing.
|
|
|
119
107
|
## Lifecycle routes (mode-specific scope identity required)
|
|
120
108
|
|
|
121
109
|
- Preparation: substantial work follows the align -> spec -> tickets ->
|
|
122
|
-
handoff path above, then
|
|
110
|
+
handoff path above, then continues under autopilot when eligible; small work uses the driver-captured
|
|
123
111
|
small-change intent.
|
|
124
112
|
- Execution: with its identity present, `axstack-implement` ->
|
|
125
113
|
`axstack-review` -> `axstack-watch`.
|
|
@@ -151,6 +151,8 @@ Authority: <who authorized which mutation>
|
|
|
151
151
|
Intent: <approved spec rev | small-change intent | adopted snapshot | peer/read-only mode>
|
|
152
152
|
Routing: <preset + source + snapshot ref>
|
|
153
153
|
Notification policy: <none | transport/target label/host/instructions path>
|
|
154
|
+
Autopilot: on | paused (<hold>; resume: <condition>) | off (cancelled <ts>); next: <step>
|
|
155
|
+
Release: <AGENTS.md file:line + tag-triggered workflow path + named install hosts> | not applicable (<reason>)
|
|
154
156
|
Source base: <exact revision or source identity>
|
|
155
157
|
IDs: <repo/project + workspace/agent receipt pointers>
|
|
156
158
|
Worktrees in other repositories: <per-run repository and worktree IDs or none>
|
|
@@ -0,0 +1,46 @@
|
|
|
1
|
+
import { readFileSync } from 'node:fs';
|
|
2
|
+
|
|
3
|
+
const args = process.argv.slice(2);
|
|
4
|
+
const option = (name) => {
|
|
5
|
+
const index = args.indexOf(name);
|
|
6
|
+
return index < 0 ? null : args[index + 1];
|
|
7
|
+
};
|
|
8
|
+
const path = option('--catalog');
|
|
9
|
+
const modelClass = option('--class');
|
|
10
|
+
const effort = option('--effort');
|
|
11
|
+
|
|
12
|
+
try {
|
|
13
|
+
if (!path || !['astra', 'sol', 'luna'].includes(modelClass) || !effort) {
|
|
14
|
+
throw new Error('expected --catalog path --class astra|sol|luna --effort level');
|
|
15
|
+
}
|
|
16
|
+
const catalog = JSON.parse(readFileSync(path, 'utf8'));
|
|
17
|
+
if (!Array.isArray(catalog.models) || typeof catalog.client_version !== 'string'
|
|
18
|
+
|| typeof catalog.fetched_at !== 'string') {
|
|
19
|
+
throw new Error('malformed catalog');
|
|
20
|
+
}
|
|
21
|
+
const classPattern = new RegExp(`^gpt-(\\d+(?:\\.\\d+)*)-${modelClass}$`);
|
|
22
|
+
const candidates = catalog.models
|
|
23
|
+
.filter((entry) => entry && entry.visibility === 'list'
|
|
24
|
+
&& typeof entry.slug === 'string'
|
|
25
|
+
&& classPattern.test(entry.slug)
|
|
26
|
+
&& Array.isArray(entry.supported_reasoning_levels)
|
|
27
|
+
&& entry.supported_reasoning_levels.some((level) => level?.effort === effort))
|
|
28
|
+
.map((entry) => entry.slug)
|
|
29
|
+
.sort((left, right) => {
|
|
30
|
+
const a = left.slice(4, -(modelClass.length + 1)).split('.').map(Number);
|
|
31
|
+
const b = right.slice(4, -(modelClass.length + 1)).split('.').map(Number);
|
|
32
|
+
for (let i = 0; i < Math.max(a.length, b.length); i++) {
|
|
33
|
+
const difference = (b[i] ?? 0) - (a[i] ?? 0);
|
|
34
|
+
if (difference) return difference;
|
|
35
|
+
}
|
|
36
|
+
return 0;
|
|
37
|
+
});
|
|
38
|
+
if (!candidates.length) throw new Error(`no eligible ${modelClass} model in catalog`);
|
|
39
|
+
console.log(JSON.stringify({
|
|
40
|
+
provider: 'codex', modelClass, effort, model: candidates[0], candidates,
|
|
41
|
+
catalog: { path, client_version: catalog.client_version, fetched_at: catalog.fetched_at },
|
|
42
|
+
}));
|
|
43
|
+
} catch (error) {
|
|
44
|
+
console.error(`model catalog resolution hold: ${error.message}`);
|
|
45
|
+
process.exitCode = 1;
|
|
46
|
+
}
|
|
@@ -5,6 +5,9 @@ description: When exploring or planning engineering work, use axstack-align to s
|
|
|
5
5
|
|
|
6
6
|
# Align
|
|
7
7
|
|
|
8
|
+
For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
|
|
9
|
+
for phase continuation and holds.
|
|
10
|
+
|
|
8
11
|
On driver entry, sweep under [Workspace hygiene](../axstack/references/workspace-hygiene.md); dispatched workers do not sweep.
|
|
9
12
|
For every dispatch brief, name its private `<run dir>/evidence/<dispatch>/` folder.
|
|
10
13
|
|
|
@@ -166,7 +169,7 @@ record or spec. Read-only scope keeps proposed documentation in the permitted
|
|
|
166
169
|
private record or response. Documentation is neither implementation nor spec
|
|
167
170
|
approval; record chosen document names and paths once per run.
|
|
168
171
|
|
|
169
|
-
## Read back, classify, and
|
|
172
|
+
## Read back, classify, and route
|
|
170
173
|
|
|
171
174
|
1. Read back the decisions, constraints, exclusions, and remaining evidence
|
|
172
175
|
gaps. For substantial work, this summary becomes part of the draft spec in
|
|
@@ -189,6 +192,7 @@ approval; record chosen document names and paths once per run.
|
|
|
189
192
|
[Orca runtime](../axstack/references/orca-runtime.md) immediately before
|
|
190
193
|
actual dispatch. Alignment completion never dispatches a recipient.
|
|
191
194
|
|
|
192
|
-
Alignment
|
|
193
|
-
identity is explicit
|
|
194
|
-
|
|
195
|
+
Alignment completes for both sizes only when the handoff is usable and its next
|
|
196
|
+
scope identity is explicit. An eligible delivery run continues under Autopilot;
|
|
197
|
+
an explicit stop-after-Align request ends here. Substantial work continues to
|
|
198
|
+
Spec, and small work continues from its small-change intent to Implement.
|
|
@@ -28,8 +28,13 @@ immediately before an actual auditor profile or session dispatch. Ordinary
|
|
|
28
28
|
audit reading and record writing do not load it, and the auditor never
|
|
29
29
|
dispatches.
|
|
30
30
|
|
|
31
|
-
Core owns the `axstack-auditor` profile (
|
|
32
|
-
|
|
31
|
+
Core owns the `axstack-auditor` profile (claude/sonnet high in
|
|
32
|
+
mixed/claude-only; codex/luna xhigh in codex-only) and its invocation.
|
|
33
|
+
This skill governs what that auditor reads, measures, and proposes.
|
|
34
|
+
Dispatch `axstack-auditor` and `axstack-auditor-sol` independently on the same
|
|
35
|
+
bounded brief, without cross-reading. The driver reconciles findings per claim;
|
|
36
|
+
never average verdicts. Record an intentionally absent Sol seat and continue
|
|
37
|
+
with the base auditor alone; a configured but unavailable seat holds its work.
|
|
33
38
|
The user-chosen improvement mode is a tested, independently reviewed PR that a
|
|
34
39
|
human merges.
|
|
35
40
|
|
|
@@ -86,6 +91,14 @@ counts with denominators plus the evidence behind the count:
|
|
|
86
91
|
evidence path is absent, or `UNKNOWN` with the reason when its records are
|
|
87
92
|
unavailable.
|
|
88
93
|
- Independent exact-revision review status and unresolved findings.
|
|
94
|
+
For authored PRs, measure the selected reviewer against this class pairing:
|
|
95
|
+
|
|
96
|
+
| Preset | Author class | Reviewer (class/effort) |
|
|
97
|
+
| --- | --- | --- |
|
|
98
|
+
| `mixed` | `codex/sol` | `axstack-reviewer-secondary` (`claude/opus` medium) |
|
|
99
|
+
| `mixed` | `claude/opus` | `axstack-reviewer-primary` (`codex/sol` high) |
|
|
100
|
+
| `codex-only` | `codex/sol` | `axstack-reviewer-secondary` (`codex/luna` xhigh) |
|
|
101
|
+
| `claude-only` | `claude/opus` | `axstack-reviewer-secondary` (`claude/sonnet` high) |
|
|
89
102
|
- Simplification applicability determinations evidenced / total candidate diffs,
|
|
90
103
|
and complete simplification receipts / total candidates, broken down as
|
|
91
104
|
`applied`, `not-applicable`, or `UNKNOWN` with the reason. This measures
|
|
@@ -5,6 +5,9 @@ description: When an approved task is ready to build or repair, use axstack-impl
|
|
|
5
5
|
|
|
6
6
|
# Implement
|
|
7
7
|
|
|
8
|
+
For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
|
|
9
|
+
for phase continuation and holds.
|
|
10
|
+
|
|
8
11
|
On driver entry, sweep under [Workspace hygiene](../axstack/references/workspace-hygiene.md); dispatched workers do not sweep.
|
|
9
12
|
For every dispatch brief, name its private `<run dir>/evidence/<dispatch>/` folder.
|
|
10
13
|
Include [Safe deletion](../axstack/references/workspace-hygiene.md#safe-deletion) in author briefs.
|
|
@@ -206,27 +209,43 @@ For each PR:
|
|
|
206
209
|
`REQUEST_CHANGES`, a failed required check, or post-readiness feedback returns
|
|
207
210
|
findings to the same author for a new revision, increments `repairs`, and
|
|
208
211
|
returns to step 1. `INCOMPLETE`, a provenance gap, unavailable model, serious
|
|
209
|
-
risk, or the third
|
|
210
|
-
parent sends its child back
|
|
212
|
+
risk, or the third review round with `REQUEST_CHANGES` and/or diligence
|
|
213
|
+
`FINDINGS` on one PR records `held`. A changed parent sends its child back
|
|
214
|
+
to step 1.
|
|
215
|
+
Merge-ready also requires a current diligence `PASS` at that head; diligence
|
|
216
|
+
`FINDINGS` return to the same author within the review round.
|
|
217
|
+
A round with reviewer `REQUEST_CHANGES` and/or diligence `FINDINGS` increments
|
|
218
|
+
`repairs` once and counts once toward the third-round hold.
|
|
211
219
|
|
|
212
220
|
One run-level completion wait covers every unsettled Dispatch; the bounded
|
|
213
|
-
forge check wait is the only other wait.
|
|
221
|
+
forge check wait is the only other implementation wait. The eligible run arms
|
|
222
|
+
one maintain-mode chat-run watch at its first published PR; that watch owns its
|
|
223
|
+
10-minute harness wake or Orca fallback. End a turn only when every required PR
|
|
214
224
|
is `merge-ready` or `held`. Under the recorded Notification policy,
|
|
215
|
-
`axstack-relay` sends only a serious risk immediately
|
|
216
|
-
operation
|
|
217
|
-
|
|
225
|
+
`axstack-relay` sends only a serious risk immediately, a genuine blocked
|
|
226
|
+
operation needing user intervention after bounded safe recovery, or the
|
|
227
|
+
decision holds and capped milestones named by the recorded Notification policy.
|
|
228
|
+
Routine questions stay in Orca. Progress, CI pending, and completion always stay
|
|
218
229
|
in Orca.
|
|
230
|
+
Only the bounded categories—user-decision holds (including spec approval),
|
|
231
|
+
serious-risk holds, and at most two merge-ready/merged milestones per run—may
|
|
232
|
+
be relayed under the recorded Notification policy.
|
|
219
233
|
|
|
220
234
|
Merge-ready is the human boundary: the user merges, bottom-up for a stack. The
|
|
221
|
-
driver resumes on the user's next message
|
|
222
|
-
|
|
235
|
+
driver resumes on the user's next message, `/axstack-watch`, or the armed
|
|
236
|
+
chat-run watch wake; no Orca merge wake exists today.
|
|
237
|
+
Re-read forge state: record forge-merged PRs as `merged`;
|
|
223
238
|
changed heads or feedback return to step 1; retain useful author work before Close-out.
|
|
224
|
-
Run Close-out once only after every required PR is forge-merged
|
|
225
|
-
|
|
239
|
+
Run Close-out once only after every required PR is forge-merged, the run's
|
|
240
|
+
Release step is settled or not applicable, and acceptance passes. It settles
|
|
241
|
+
workers, records counts, makes the auditor decision and
|
|
226
242
|
settlement, releases worktrees, closes eligible tickets, and archives the run.
|
|
243
|
+
Without an Autopilot or Release record, the Release step is not applicable for
|
|
244
|
+
both Close-out and run completion.
|
|
227
245
|
|
|
228
246
|
The loop requires the `mixed` two-provider authored-review row. `codex-only` or
|
|
229
247
|
`claude-only` holds at step (3) for an explicit user routing choice, with no
|
|
230
248
|
substitution or same-provider review. Derived PR states are `authoring |
|
|
231
249
|
published | in-review | repairing(n) | merge-ready | merged | held`. The run is
|
|
232
|
-
done only when every required PR is forge-merged
|
|
250
|
+
done only when every required PR is forge-merged, the run's Release step is
|
|
251
|
+
settled or not applicable, and Close-out has receipts.
|
|
@@ -19,6 +19,11 @@ role dispatch, load the [Orca runtime
|
|
|
19
19
|
sequence](../axstack/references/orca-runtime.md). Use existing
|
|
20
20
|
`axstack-explore-codebase` or `axstack-research-code` roles only when their
|
|
21
21
|
specialization materially helps; create no new profile.
|
|
22
|
+
When dispatching `axstack-research-code`, dispatch `axstack-research-code-sol`
|
|
23
|
+
independently on the same bounded brief without cross-reading. The driver
|
|
24
|
+
reconciles findings per claim and never averages them. Record an intentionally
|
|
25
|
+
absent Sol pair and proceed with the base seat alone; a configured but
|
|
26
|
+
unavailable pair holds its work.
|
|
22
27
|
|
|
23
28
|
## 1. Bound discovery
|
|
24
29
|
|
|
@@ -5,6 +5,9 @@ description: When the user requests a relay message or test, or an authorized no
|
|
|
5
5
|
|
|
6
6
|
# Relay
|
|
7
7
|
|
|
8
|
+
For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
|
|
9
|
+
for phase continuation and holds.
|
|
10
|
+
|
|
8
11
|
Send normal messages, transport tests, and authorized notifications to the
|
|
9
12
|
user through Hermes' native one-way `hermes send`. This is an inline caller
|
|
10
13
|
procedure: it creates no driver, team, owner, auditor, monitor, child session,
|
|
@@ -25,8 +28,12 @@ Choose the applicable message type:
|
|
|
25
28
|
in the caller's private notification policy. State the issue, impact, and the
|
|
26
29
|
answer or action needed.
|
|
27
30
|
- **Routine run events:** questions, spec approvals, progress, CI pending,
|
|
28
|
-
merge-ready, merged, and completion stay in Orca
|
|
29
|
-
|
|
31
|
+
merge-ready, merged, and completion stay in Orca unless the recorded
|
|
32
|
+
Notification policy names it. A policy may name only user-decision holds and
|
|
33
|
+
at most two merge-ready/merged milestones per run; deduplicate across implementation
|
|
34
|
+
and release. Progress, CI pending, and completion are never eligible merely
|
|
35
|
+
because a policy exists. They never become proactive relay messages merely
|
|
36
|
+
because the run is waiting.
|
|
30
37
|
|
|
31
38
|
Verify the transport, execution host, and intended recipient from the user's
|
|
32
39
|
request, trusted caller context, or an existing private notification policy.
|
|
@@ -29,8 +29,9 @@ is part of research.
|
|
|
29
29
|
|
|
30
30
|
2. **Fan out research:** A single factual lookup stays in the current chat.
|
|
31
31
|
Every other research run dispatches every configured research branch through
|
|
32
|
-
Orca: requirements
|
|
33
|
-
(Gemini/Antigravity, with Google Search
|
|
32
|
+
Orca: requirements, code, and web (Sonnet high in mixed/claude-only;
|
|
33
|
+
Codex in codex-only), web-google (Gemini/Antigravity, with Google Search
|
|
34
|
+
built in), and X (Grok).
|
|
34
35
|
Give each branch one owner, allow no cross-reading, and require a cited note
|
|
35
36
|
with a URL and access date per claim; re-open sources and never trust a search
|
|
36
37
|
summary. The driver reconciles agreements/disagreements per claim.
|
|
@@ -41,16 +42,27 @@ is part of research.
|
|
|
41
42
|
|
|
42
43
|
- `axstack-research-requirements`: requirements and intent.
|
|
43
44
|
- `axstack-research-code`: code behavior.
|
|
45
|
+
- `axstack-research-code-sol`: independent Sol code investigation.
|
|
44
46
|
- `axstack-research-web`: web and external sources.
|
|
45
47
|
- `axstack-research-web-google`: Google-Search-grounded web sources via Gemini/Antigravity.
|
|
46
48
|
- `axstack-research-x`: X (Twitter) posts and threads via Grok; cite post URLs and dates.
|
|
47
49
|
- `axstack-explore-codebase`: broad codebase mapping.
|
|
48
50
|
- `axstack-explore-execution`: execution and runtime traces.
|
|
51
|
+
- `axstack-explore-execution-sol`: independent Sol execution investigation.
|
|
52
|
+
|
|
53
|
+
When dispatching `axstack-research-code` or `axstack-explore-execution`,
|
|
54
|
+
dispatch its `-sol` pair independently on the same bounded brief without
|
|
55
|
+
cross-reading. The driver reconciles agreement and disagreement per claim,
|
|
56
|
+
never averaging findings. Record an intentionally absent pair and proceed
|
|
57
|
+
with the base seat alone; a configured but unavailable seat holds its work.
|
|
49
58
|
|
|
50
59
|
3. **Gather primary source evidence.** Inspect the actual documentation, code,
|
|
51
60
|
or tool output for every answer-changing claim. Apply the source standards
|
|
52
61
|
for citations, freshness, revisions, and access dates. Continue until each
|
|
53
62
|
material claim has direct evidence or a named evidence gap.
|
|
63
|
+
Before the driver folds verified claims, dispatch `axstack-diligence` under
|
|
64
|
+
[Diligence](../axstack/references/diligence.md) to reopen cited sources for
|
|
65
|
+
answer-changing claims and flag mismatches.
|
|
54
66
|
|
|
55
67
|
4. **Form the verdict.** Mark every material claim as **verified**,
|
|
56
68
|
**inference**, or **unverified** using the source standards. Derive
|