axstack 0.20.27 → 0.20.29
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +4 -1
- package/docs/installation.md +6 -2
- package/docs/workflows.md +6 -2
- package/package.json +1 -1
- package/profiles/presets/claude-only.json +29 -20
- package/profiles/presets/codex-only.json +10 -1
- package/profiles/presets/mixed.json +22 -13
- package/skills/axstack/references/orca-runtime.md +1 -1
- package/skills/axstack/references/routing.md +14 -12
- package/skills/axstack/references/ui-verification.md +13 -0
- package/skills/axstack-debug/SKILL.md +2 -0
- package/skills/axstack-explain/references/visual-qa.md +9 -6
- package/skills/axstack-implement/SKILL.md +3 -2
- package/skills/axstack-review/SKILL.md +4 -3
- package/src/roles.js +1 -1
package/README.md
CHANGED
|
@@ -100,13 +100,16 @@ upgrades, conflicts, and uninstalling.
|
|
|
100
100
|
- Agents keep accepted decisions and evidence for resume. Missing authority,
|
|
101
101
|
unavailable models, and serious risks surface as holds. The human merges by default.
|
|
102
102
|
|
|
103
|
-
Choose one explicit preset (
|
|
103
|
+
Choose one explicit preset (28 roles each): [mixed](profiles/presets/mixed.json)
|
|
104
104
|
(recommended), [codex-only](profiles/presets/codex-only.json), or
|
|
105
105
|
[claude-only](profiles/presets/claude-only.json). Mixed supports cross-provider
|
|
106
106
|
implementation review; single-provider presets have workflow limits and are
|
|
107
107
|
not automatic fallbacks when a model is unavailable. See
|
|
108
108
|
[workflow and routing details](docs/workflows.md).
|
|
109
109
|
|
|
110
|
+
In `mixed` and `claude-only`, requirements research, web research, and the
|
|
111
|
+
optional monitor use Claude Sonnet 5.5 high.
|
|
112
|
+
|
|
110
113
|
## Optional PR automation
|
|
111
114
|
|
|
112
115
|
Manual review works without a schedule. Own open PRs in chat-run mode use a
|
package/docs/installation.md
CHANGED
|
@@ -71,7 +71,7 @@ profiles/presets/codex-only.json
|
|
|
71
71
|
profiles/presets/claude-only.json
|
|
72
72
|
```
|
|
73
73
|
|
|
74
|
-
Each has exactly `{ "version": 1, "roles": [...] }` with the same
|
|
74
|
+
Each has exactly `{ "version": 1, "roles": [...] }` with the same 28 stable
|
|
75
75
|
role IDs. Installation writes `<skills-dir>/axstack/roles.json` as
|
|
76
76
|
`{ "version": 1, "preset": "<selected preset>", "roles": [...] }` and records
|
|
77
77
|
its ownership hash like every other installed skill asset. There is no second
|
|
@@ -172,10 +172,14 @@ to rewrite them.
|
|
|
172
172
|
## Role behavior after installation
|
|
173
173
|
|
|
174
174
|
The runtime reads `roles.json` from the installed shared root `skills/axstack/`.
|
|
175
|
-
A new run records the selected preset plus all
|
|
175
|
+
A new run records the selected preset plus all 28 role rows. An active run keeps
|
|
176
176
|
that snapshot after a later preset install unless the user explicitly changes
|
|
177
177
|
it and accepts the resulting evidence invalidation.
|
|
178
178
|
|
|
179
|
+
The `mixed` and `claude-only` presets assign `axstack-research-requirements`,
|
|
180
|
+
`axstack-research-web`, and `axstack-monitor` to Claude Sonnet 5.5 high.
|
|
181
|
+
The `codex-only` assignments for these roles are unchanged.
|
|
182
|
+
|
|
179
183
|
The mixed checker and `axstack-research-web-google` have provider
|
|
180
184
|
`antigravity`; mixed `axstack-research-x` has provider `grok`. All three use
|
|
181
185
|
`model: null` because Orca exposes no model override for those agent-ID routes;
|
package/docs/workflows.md
CHANGED
|
@@ -48,7 +48,7 @@ only affected work.
|
|
|
48
48
|
|
|
49
49
|
Installation requires one explicit canonical preset. The three bundle files
|
|
50
50
|
under `profiles/presets/` each contain exactly
|
|
51
|
-
`{ "version": 1, "roles": [...] }` and the same
|
|
51
|
+
`{ "version": 1, "roles": [...] }` and the same 28 stable IDs.
|
|
52
52
|
|
|
53
53
|
The current chat drives on whatever model runs it; no preset carries a driver
|
|
54
54
|
role.
|
|
@@ -57,7 +57,11 @@ role.
|
|
|
57
57
|
| --- | --- | --- | --- | --- |
|
|
58
58
|
| `mixed` | Sol high | Sol high; Opus medium | Astra high / Opus xhigh | Luna xhigh |
|
|
59
59
|
| `codex-only` | Sol high | Sol high; Luna xhigh | Astra high / unavailable | Luna xhigh |
|
|
60
|
-
| `claude-only` | Opus medium | Opus medium; Sonnet
|
|
60
|
+
| `claude-only` | Opus medium | Opus medium; Sonnet high | unavailable / Opus xhigh | Sonnet high |
|
|
61
|
+
|
|
62
|
+
In `mixed` and `claude-only`, `axstack-research-requirements`,
|
|
63
|
+
`axstack-research-web`, and `axstack-monitor` use Claude Sonnet 5.5 high.
|
|
64
|
+
`codex-only` keeps its Codex assignments for those roles.
|
|
61
65
|
|
|
62
66
|
The installed `<skills-dir>/axstack/roles.json` adds the selected preset name:
|
|
63
67
|
`{ "version": 1, "preset": "<name>", "roles": [...] }`. The runtime reads it
|
package/package.json
CHANGED
|
@@ -44,7 +44,7 @@
|
|
|
44
44
|
"model": "claude-opus-5-5",
|
|
45
45
|
"modeId": "bypassPermissions",
|
|
46
46
|
"thinkingOptionId": "medium",
|
|
47
|
-
"notes": "Primary reviewer in the ordered claude-only peer pair: Opus medium followed by Sonnet
|
|
47
|
+
"notes": "Primary reviewer in the ordered claude-only peer pair: Opus medium followed by Sonnet high. Never author or owner; review exact SHA and base across all six angles and acceptance. Peer first pass stays isolated."
|
|
48
48
|
},
|
|
49
49
|
{
|
|
50
50
|
"id": "axstack-reviewer-secondary",
|
|
@@ -52,8 +52,8 @@
|
|
|
52
52
|
"provider": "claude",
|
|
53
53
|
"model": "claude-sonnet-5-5",
|
|
54
54
|
"modeId": "bypassPermissions",
|
|
55
|
-
"thinkingOptionId": "
|
|
56
|
-
"notes": "Secondary reviewer in the ordered claude-only peer pair: Opus medium followed by Sonnet
|
|
55
|
+
"thinkingOptionId": "high",
|
|
56
|
+
"notes": "Secondary reviewer in the ordered claude-only peer pair: Opus medium followed by Sonnet high. Eligible authored reviewer for an Opus-authored candidate. Never author or owner; review exact SHA and base across all six angles and acceptance. Peer first pass stays isolated."
|
|
57
57
|
},
|
|
58
58
|
{
|
|
59
59
|
"id": "axstack-checker",
|
|
@@ -61,17 +61,17 @@
|
|
|
61
61
|
"provider": "claude",
|
|
62
62
|
"model": "claude-sonnet-5-5",
|
|
63
63
|
"modeId": "bypassPermissions",
|
|
64
|
-
"thinkingOptionId": "
|
|
64
|
+
"thinkingOptionId": "high",
|
|
65
65
|
"notes": "Report-only discrepancy checker for the selected external tracker (Linear or GitHub Issues). Never mutates the tracker; the driver independently verifies evidence before applying updates. A null model means explicit user selection is required before dispatch and must never launch a provider default."
|
|
66
66
|
},
|
|
67
67
|
{
|
|
68
68
|
"id": "axstack-research-requirements",
|
|
69
69
|
"name": "Axstack research requirements",
|
|
70
70
|
"provider": "claude",
|
|
71
|
-
"model": "claude-
|
|
71
|
+
"model": "claude-sonnet-5-5",
|
|
72
72
|
"modeId": "bypassPermissions",
|
|
73
|
-
"thinkingOptionId": "
|
|
74
|
-
"notes": "
|
|
73
|
+
"thinkingOptionId": "high",
|
|
74
|
+
"notes": "Sonnet high research requirements analyst: scopes bounded questions and acceptance for a research task. Validate configured availability at launch; hold affected work without fallback."
|
|
75
75
|
},
|
|
76
76
|
{
|
|
77
77
|
"id": "axstack-research-code",
|
|
@@ -88,8 +88,8 @@
|
|
|
88
88
|
"provider": "claude",
|
|
89
89
|
"model": "claude-sonnet-5-5",
|
|
90
90
|
"modeId": "bypassPermissions",
|
|
91
|
-
"thinkingOptionId": "
|
|
92
|
-
"notes": "
|
|
91
|
+
"thinkingOptionId": "high",
|
|
92
|
+
"notes": "Sonnet high research web reader: gathers primary-source facts efficiently. Validate configured availability at launch; hold affected work without fallback."
|
|
93
93
|
},
|
|
94
94
|
{
|
|
95
95
|
"id": "axstack-research-web-google",
|
|
@@ -115,7 +115,7 @@
|
|
|
115
115
|
"provider": "claude",
|
|
116
116
|
"model": "claude-sonnet-5-5",
|
|
117
117
|
"modeId": "bypassPermissions",
|
|
118
|
-
"thinkingOptionId": "
|
|
118
|
+
"thinkingOptionId": "high",
|
|
119
119
|
"notes": "Complex visual explanation author: traces systems, changes, and implementation gaps in requested artifacts and verifies rendered behavior where applicable. Validate configured availability at launch; hold affected work without fallback."
|
|
120
120
|
},
|
|
121
121
|
{
|
|
@@ -125,7 +125,16 @@
|
|
|
125
125
|
"model": "claude-sonnet-5-5",
|
|
126
126
|
"modeId": "bypassPermissions",
|
|
127
127
|
"thinkingOptionId": "high",
|
|
128
|
-
"notes": "Independent visual explanation reviewer: checks the exact artifact for source fidelity
|
|
128
|
+
"notes": "Independent visual explanation reviewer: checks the exact artifact for text and source fidelity. The rendered pass belongs to axstack-ui-verifier. Any artifact change invalidates its review. Validate configured availability at launch; hold affected work without fallback."
|
|
129
|
+
},
|
|
130
|
+
{
|
|
131
|
+
"id": "axstack-ui-verifier",
|
|
132
|
+
"name": "Axstack UI verifier",
|
|
133
|
+
"provider": "claude",
|
|
134
|
+
"model": "claude-sonnet-5-5",
|
|
135
|
+
"modeId": "bypassPermissions",
|
|
136
|
+
"thinkingOptionId": "high",
|
|
137
|
+
"notes": "Read-only UI verifier: runs Playwright or the browser against the given build, URL, or artifact; captures screenshots, interactions, accessibility, desktop/mobile, and reduced-motion evidence in the dispatch evidence folder; returns a verdict with evidence paths. Never edits source. Validate configured availability at launch; hold affected work without fallback."
|
|
129
138
|
},
|
|
130
139
|
{
|
|
131
140
|
"id": "axstack-explore-codebase",
|
|
@@ -133,7 +142,7 @@
|
|
|
133
142
|
"provider": "claude",
|
|
134
143
|
"model": "claude-sonnet-5-5",
|
|
135
144
|
"modeId": "bypassPermissions",
|
|
136
|
-
"thinkingOptionId": "
|
|
145
|
+
"thinkingOptionId": "high",
|
|
137
146
|
"notes": "Codebase mapper: explores repository structure and interfaces for research and handoff context. Validate configured availability at launch; hold affected work without fallback."
|
|
138
147
|
},
|
|
139
148
|
{
|
|
@@ -142,7 +151,7 @@
|
|
|
142
151
|
"provider": "claude",
|
|
143
152
|
"model": "claude-sonnet-5-5",
|
|
144
153
|
"modeId": "bypassPermissions",
|
|
145
|
-
"thinkingOptionId": "
|
|
154
|
+
"thinkingOptionId": "high",
|
|
146
155
|
"notes": "Execution explorer: runs bounded checks of runtime behavior where authorized. Validate configured availability at launch; hold affected work without fallback."
|
|
147
156
|
},
|
|
148
157
|
{
|
|
@@ -151,8 +160,8 @@
|
|
|
151
160
|
"provider": "claude",
|
|
152
161
|
"model": "claude-sonnet-5-5",
|
|
153
162
|
"modeId": "bypassPermissions",
|
|
154
|
-
"thinkingOptionId": "
|
|
155
|
-
"notes": "Optional independent read-only observer for a standalone PR watch. Reads GitHub, feedback, and checks, persists event IDs, and wakes the owner only for a new actionable event. Never sends, authors, reviews, replies, or acts as either reusable PR manager. Healthy snapshots stay quiet. Chat-run mode: one same-host native read-only observer per Run; fresh finite passes report precise deltas internally to the original Run/driver and may disable/read back only their own automation at verified stop. No repair, dispatch, public notification, or replacement coordinator. Effective scheduled model/effort and wake require live proof."
|
|
163
|
+
"thinkingOptionId": "high",
|
|
164
|
+
"notes": "Optional Sonnet high independent read-only observer for a standalone PR watch. Reads GitHub, feedback, and checks, persists event IDs, and wakes the owner only for a new actionable event. Never sends, authors, reviews, replies, or acts as either reusable PR manager. Healthy snapshots stay quiet. Chat-run mode: one same-host native read-only observer per Run; fresh finite passes report precise deltas internally to the original Run/driver and may disable/read back only their own automation at verified stop. No repair, dispatch, public notification, or replacement coordinator. Effective scheduled model/effort and wake require live proof."
|
|
156
165
|
},
|
|
157
166
|
{
|
|
158
167
|
"id": "axstack-auditor",
|
|
@@ -160,7 +169,7 @@
|
|
|
160
169
|
"provider": "claude",
|
|
161
170
|
"model": "claude-sonnet-5-5",
|
|
162
171
|
"modeId": "bypassPermissions",
|
|
163
|
-
"thinkingOptionId": "
|
|
172
|
+
"thinkingOptionId": "high",
|
|
164
173
|
"notes": "Read-only end-of-run and checkpoint auditor. Collects scope and outcome evidence with counts and denominators and reports PASS, FAIL, or UNKNOWN without inventing numbers. Never edits, merges, activates, or audits itself."
|
|
165
174
|
},
|
|
166
175
|
{
|
|
@@ -178,8 +187,8 @@
|
|
|
178
187
|
"provider": "claude",
|
|
179
188
|
"model": "claude-sonnet-5-5",
|
|
180
189
|
"modeId": "bypassPermissions",
|
|
181
|
-
"thinkingOptionId": "
|
|
182
|
-
"notes": "Debug investigator seat 2. Dispatched only by axstack-debug at L1 with the shared evidence packet and one distinct brief; never reads another investigator's output. Works in its own disposable worktree at the pinned revision plus the recorded dirty patch; may instrument there for probes; never commits, pushes, publishes, or creates children. Returns one receipt per brief. Independence comes from brief isolation, not model diversity. This preset repeats claude-sonnet-5-5 at
|
|
190
|
+
"thinkingOptionId": "high",
|
|
191
|
+
"notes": "Debug investigator seat 2. Dispatched only by axstack-debug at L1 with the shared evidence packet and one distinct brief; never reads another investigator's output. Works in its own disposable worktree at the pinned revision plus the recorded dirty patch; may instrument there for probes; never commits, pushes, publishes, or creates children. Returns one receipt per brief. Independence comes from brief isolation, not model diversity. This preset repeats claude-sonnet-5-5 at high effort because it has fewer model families; independence comes from brief isolation."
|
|
183
192
|
},
|
|
184
193
|
{
|
|
185
194
|
"id": "axstack-debug-investigator-3",
|
|
@@ -196,8 +205,8 @@
|
|
|
196
205
|
"provider": "claude",
|
|
197
206
|
"model": "claude-sonnet-5-5",
|
|
198
207
|
"modeId": "bypassPermissions",
|
|
199
|
-
"thinkingOptionId": "
|
|
200
|
-
"notes": "Debug investigator seat 4. Dispatched only by axstack-debug at L1 with the shared evidence packet and one distinct brief; never reads another investigator's output. Works in its own disposable worktree at the pinned revision plus the recorded dirty patch; may instrument there for probes; never commits, pushes, publishes, or creates children. Returns one receipt per brief. Independence comes from brief isolation, not model diversity. This preset repeats claude-sonnet-5-5 at
|
|
208
|
+
"thinkingOptionId": "high",
|
|
209
|
+
"notes": "Debug investigator seat 4. Dispatched only by axstack-debug at L1 with the shared evidence packet and one distinct brief; never reads another investigator's output. Works in its own disposable worktree at the pinned revision plus the recorded dirty patch; may instrument there for probes; never commits, pushes, publishes, or creates children. Returns one receipt per brief. Independence comes from brief isolation, not model diversity. This preset repeats claude-sonnet-5-5 at high effort because it has fewer model families; independence comes from brief isolation."
|
|
201
210
|
},
|
|
202
211
|
{
|
|
203
212
|
"id": "axstack-arena-judge-astra",
|
|
@@ -125,7 +125,16 @@
|
|
|
125
125
|
"model": "gpt-6-luna",
|
|
126
126
|
"modeId": "full-access",
|
|
127
127
|
"thinkingOptionId": "xhigh",
|
|
128
|
-
"notes": "Independent visual explanation reviewer: checks the exact artifact for source fidelity
|
|
128
|
+
"notes": "Independent visual explanation reviewer: checks the exact artifact for text and source fidelity. The rendered pass belongs to axstack-ui-verifier. Any artifact change invalidates its review. Validate configured availability at launch; hold affected work without fallback."
|
|
129
|
+
},
|
|
130
|
+
{
|
|
131
|
+
"id": "axstack-ui-verifier",
|
|
132
|
+
"name": "Axstack UI verifier",
|
|
133
|
+
"provider": "codex",
|
|
134
|
+
"model": "gpt-6-sol",
|
|
135
|
+
"modeId": "full-access",
|
|
136
|
+
"thinkingOptionId": "medium",
|
|
137
|
+
"notes": "Read-only UI verifier: runs Playwright or the browser against the given build, URL, or artifact; captures screenshots, interactions, accessibility, desktop/mobile, and reduced-motion evidence in the dispatch evidence folder; returns a verdict with evidence paths. Never edits source. Validate configured availability at launch; hold affected work without fallback."
|
|
129
138
|
},
|
|
130
139
|
{
|
|
131
140
|
"id": "axstack-explore-codebase",
|
|
@@ -68,10 +68,10 @@
|
|
|
68
68
|
"id": "axstack-research-requirements",
|
|
69
69
|
"name": "Axstack research requirements",
|
|
70
70
|
"provider": "claude",
|
|
71
|
-
"model": "claude-
|
|
71
|
+
"model": "claude-sonnet-5-5",
|
|
72
72
|
"modeId": "bypassPermissions",
|
|
73
|
-
"thinkingOptionId": "
|
|
74
|
-
"notes": "
|
|
73
|
+
"thinkingOptionId": "high",
|
|
74
|
+
"notes": "Sonnet high research requirements analyst: scopes bounded questions and acceptance for a research task. Validate configured availability at launch; hold affected work without fallback."
|
|
75
75
|
},
|
|
76
76
|
{
|
|
77
77
|
"id": "axstack-research-code",
|
|
@@ -86,10 +86,10 @@
|
|
|
86
86
|
"id": "axstack-research-web",
|
|
87
87
|
"name": "Axstack research web",
|
|
88
88
|
"provider": "claude",
|
|
89
|
-
"model": "claude-
|
|
89
|
+
"model": "claude-sonnet-5-5",
|
|
90
90
|
"modeId": "bypassPermissions",
|
|
91
|
-
"thinkingOptionId": "
|
|
92
|
-
"notes": "
|
|
91
|
+
"thinkingOptionId": "high",
|
|
92
|
+
"notes": "Sonnet high research web reader: gathers primary-source facts efficiently. Validate configured availability at launch; hold affected work without fallback."
|
|
93
93
|
},
|
|
94
94
|
{
|
|
95
95
|
"id": "axstack-research-web-google",
|
|
@@ -115,7 +115,7 @@
|
|
|
115
115
|
"provider": "claude",
|
|
116
116
|
"model": "claude-sonnet-5-5",
|
|
117
117
|
"modeId": "bypassPermissions",
|
|
118
|
-
"thinkingOptionId": "
|
|
118
|
+
"thinkingOptionId": "high",
|
|
119
119
|
"notes": "Complex visual explanation author: traces systems, changes, and implementation gaps in requested artifacts and verifies rendered behavior where applicable. Validate configured availability at launch; hold affected work without fallback."
|
|
120
120
|
},
|
|
121
121
|
{
|
|
@@ -125,7 +125,16 @@
|
|
|
125
125
|
"model": "gpt-6-luna",
|
|
126
126
|
"modeId": "full-access",
|
|
127
127
|
"thinkingOptionId": "xhigh",
|
|
128
|
-
"notes": "Independent visual explanation reviewer: checks the exact artifact for source fidelity
|
|
128
|
+
"notes": "Independent visual explanation reviewer: checks the exact artifact for text and source fidelity. The rendered pass belongs to axstack-ui-verifier. Any artifact change invalidates its review. Validate configured availability at launch; hold affected work without fallback."
|
|
129
|
+
},
|
|
130
|
+
{
|
|
131
|
+
"id": "axstack-ui-verifier",
|
|
132
|
+
"name": "Axstack UI verifier",
|
|
133
|
+
"provider": "claude",
|
|
134
|
+
"model": "claude-sonnet-5-5",
|
|
135
|
+
"modeId": "bypassPermissions",
|
|
136
|
+
"thinkingOptionId": "high",
|
|
137
|
+
"notes": "Read-only UI verifier: runs Playwright or the browser against the given build, URL, or artifact; captures screenshots, interactions, accessibility, desktop/mobile, and reduced-motion evidence in the dispatch evidence folder; returns a verdict with evidence paths. Never edits source. Validate configured availability at launch; hold affected work without fallback."
|
|
129
138
|
},
|
|
130
139
|
{
|
|
131
140
|
"id": "axstack-explore-codebase",
|
|
@@ -133,7 +142,7 @@
|
|
|
133
142
|
"provider": "claude",
|
|
134
143
|
"model": "claude-sonnet-5-5",
|
|
135
144
|
"modeId": "bypassPermissions",
|
|
136
|
-
"thinkingOptionId": "
|
|
145
|
+
"thinkingOptionId": "high",
|
|
137
146
|
"notes": "Codebase mapper: explores repository structure and interfaces for research and handoff context. Validate configured availability at launch; hold affected work without fallback."
|
|
138
147
|
},
|
|
139
148
|
{
|
|
@@ -149,10 +158,10 @@
|
|
|
149
158
|
"id": "axstack-monitor",
|
|
150
159
|
"name": "Axstack monitor",
|
|
151
160
|
"provider": "claude",
|
|
152
|
-
"model": "claude-
|
|
161
|
+
"model": "claude-sonnet-5-5",
|
|
153
162
|
"modeId": "bypassPermissions",
|
|
154
|
-
"thinkingOptionId": "
|
|
155
|
-
"notes": "Optional independent read-only observer for a standalone PR watch. Reads GitHub, feedback, and checks, persists event IDs, and wakes the owner only for a new actionable event. Never sends, authors, reviews, replies, or acts as either reusable PR manager. Healthy snapshots stay quiet. Chat-run mode: one same-host native read-only observer per Run; fresh finite passes report precise deltas internally to the original Run/driver and may disable/read back only their own automation at verified stop. No repair, dispatch, public notification, or replacement coordinator. Effective scheduled model/effort and wake require live proof."
|
|
163
|
+
"thinkingOptionId": "high",
|
|
164
|
+
"notes": "Optional Sonnet high independent read-only observer for a standalone PR watch. Reads GitHub, feedback, and checks, persists event IDs, and wakes the owner only for a new actionable event. Never sends, authors, reviews, replies, or acts as either reusable PR manager. Healthy snapshots stay quiet. Chat-run mode: one same-host native read-only observer per Run; fresh finite passes report precise deltas internally to the original Run/driver and may disable/read back only their own automation at verified stop. No repair, dispatch, public notification, or replacement coordinator. Effective scheduled model/effort and wake require live proof."
|
|
156
165
|
},
|
|
157
166
|
{
|
|
158
167
|
"id": "axstack-auditor",
|
|
@@ -187,7 +196,7 @@
|
|
|
187
196
|
"provider": "claude",
|
|
188
197
|
"model": "claude-sonnet-5-5",
|
|
189
198
|
"modeId": "bypassPermissions",
|
|
190
|
-
"thinkingOptionId": "
|
|
199
|
+
"thinkingOptionId": "high",
|
|
191
200
|
"notes": "Debug investigator seat 3. Dispatched only by axstack-debug at L1 with the shared evidence packet and one distinct brief; never reads another investigator's output. Works in its own disposable worktree at the pinned revision plus the recorded dirty patch; may instrument there for probes; never commits, pushes, publishes, or creates children. Returns one receipt per brief. Independence comes from brief isolation, not model diversity."
|
|
192
201
|
},
|
|
193
202
|
{
|
|
@@ -38,7 +38,7 @@ Read `roles.json` from the installed shared root `skills/axstack/`. The installe
|
|
|
38
38
|
shape is `{ "version": 1, "preset": "<name>", "roles": [...] }`. Bundled
|
|
39
39
|
profiles are setup inputs shaped as
|
|
40
40
|
`{ "version": 1, "roles": [...] }`. A new run records the selected preset and
|
|
41
|
-
all
|
|
41
|
+
all 28 role rows once. An active run keeps the exact snapshot until the user
|
|
42
42
|
explicitly changes it.
|
|
43
43
|
|
|
44
44
|
Select the requested role by stable ID. A missing or null model holds only that role;
|
|
@@ -11,10 +11,10 @@ skills root, or an explicit user selection in the run record. Missing or contrad
|
|
|
11
11
|
a setup gap: hold. Never infer from live profiles or `list_profiles`, harness,
|
|
12
12
|
tools, credentials, quota, subscription, or default to `mixed`.
|
|
13
13
|
|
|
14
|
-
At start, snapshot all
|
|
14
|
+
At start, snapshot all 28 role IDs with provider/model/mode/effort; absent
|
|
15
15
|
or unconfigured roles are recorded explicitly; never default.
|
|
16
|
-
Such a role holds only
|
|
17
|
-
|
|
16
|
+
Such a role holds only its work. Later installed or changed roles need an
|
|
17
|
+
explicit user decision to enter the snapshot. Live profiles
|
|
18
18
|
are authoritative at snapshot time and for availability; bundled presets are setup
|
|
19
19
|
inputs, not runtime proof.
|
|
20
20
|
|
|
@@ -34,7 +34,7 @@ provider/model/effort substitution.
|
|
|
34
34
|
| `mixed` | Codex / Sol (`codex/gpt-6-sol`) | `axstack-reviewer-secondary` (`claude/claude-opus-5-5` medium) |
|
|
35
35
|
| `mixed` | Claude / Opus (`claude/claude-opus-5-5`) | `axstack-reviewer-primary` (`codex/gpt-6-sol` high) |
|
|
36
36
|
| `codex-only` | Codex / Sol (`codex/gpt-6-sol`) | `axstack-reviewer-secondary` (`codex/gpt-6-luna` xhigh) |
|
|
37
|
-
| `claude-only` | Claude / Opus (`claude/claude-opus-5-5`) | `axstack-reviewer-secondary` (`claude/claude-sonnet-5-5`
|
|
37
|
+
| `claude-only` | Claude / Opus (`claude/claude-opus-5-5`) | `axstack-reviewer-secondary` (`claude/claude-sonnet-5-5` high) |
|
|
38
38
|
- `axstack-advisor-astra`/`axstack-advisor-opus` advise and author candidates;
|
|
39
39
|
`axstack-arena-candidate-grok`/
|
|
40
40
|
`axstack-arena-candidate-antigravity` add families.
|
|
@@ -42,6 +42,9 @@ provider/model/effort substitution.
|
|
|
42
42
|
High-stakes/trigger: fresh [contract](contracts.md) session.
|
|
43
43
|
`axstack-auditor` audits; `axstack-checker` reports discrepancies.
|
|
44
44
|
- `axstack-explainer`/`axstack-explainer-review`: explain/review.
|
|
45
|
+
- `axstack-ui-verifier`: [UI checks](ui-verification.md).
|
|
46
|
+
- `axstack-research-requirements`/`axstack-research-web`/`axstack-monitor`:
|
|
47
|
+
Sonnet 5.5 high in mixed/claude-only.
|
|
45
48
|
`axstack-monitor`: standalone watch never sends; chat-run watch: bounded
|
|
46
49
|
internal reports to its Run and original driver.
|
|
47
50
|
- `axstack-debug-investigator-1..4` probe L1 briefs.
|
|
@@ -55,18 +58,17 @@ step (3) for user routing: no substitution or same-provider review.
|
|
|
55
58
|
|
|
56
59
|
## Direct routes (no spec ceremony)
|
|
57
60
|
|
|
58
|
-
-
|
|
59
|
-
|
|
60
|
-
-
|
|
61
|
-
current/intended behavior
|
|
62
|
-
|
|
61
|
+
- Bounded research -> `axstack-research`: verify primary sources and code,
|
|
62
|
+
cite limits, and fan out distinct questions.
|
|
63
|
+
- Understand a system or gap -> `axstack-explain`:
|
|
64
|
+
current/intended behavior and bounded gaps from docs and renders.
|
|
65
|
+
"What could this break" follows
|
|
63
66
|
[Blast radius](blast-radius.md). Publication needs separate authority.
|
|
64
67
|
- A bug, failing test, regression, or wrong behavior, red loop wanted ->
|
|
65
68
|
`axstack-debug`: diagnose, escalate via adviser-directed investigators, hand
|
|
66
69
|
off a classified repair (explain: how; debug: what's wrong).
|
|
67
|
-
-
|
|
68
|
-
|
|
69
|
-
edits.
|
|
70
|
+
- Code quality/refactor discovery -> `axstack-improve`: rank bounded
|
|
71
|
+
candidates with evidence; report only, no source edits.
|
|
70
72
|
- Accepted worker/Task/Run completion or bounded backlog request -> driver invokes
|
|
71
73
|
`axstack-cleanup` inline; never dispatch it.
|
|
72
74
|
- Preparation completion, watch expiry, resume, or reconciliation -> the
|
|
@@ -0,0 +1,13 @@
|
|
|
1
|
+
# UI verification
|
|
2
|
+
|
|
3
|
+
Every Playwright, browser, or rendered-UI check, including a "confirm it in the
|
|
4
|
+
browser" step, goes through an Orca Dispatch to `axstack-ui-verifier` from the
|
|
5
|
+
run's role snapshot. Give it the exact build, URL, or artifact and the private
|
|
6
|
+
dispatch's evidence folder. The verifier is read-only: it never edits source.
|
|
7
|
+
The PR writer remains the sole writer.
|
|
8
|
+
|
|
9
|
+
Ask for screenshots and observed interactions, accessibility, desktop and
|
|
10
|
+
mobile layouts, and reduced-motion behavior where relevant. The verifier
|
|
11
|
+
returns a verdict with evidence paths and names checks it could not run.
|
|
12
|
+
Keep the verdict tied to the exact artifact or revision; changed bytes need a
|
|
13
|
+
fresh rendered pass.
|
|
@@ -39,6 +39,8 @@ recorded reason.
|
|
|
39
39
|
loop cannot be built, stop, list what was tried, and ask the user for an
|
|
40
40
|
environment, a redacted artifact, or instrumentation permission. Done when
|
|
41
41
|
the command has run once and its red output is recorded.
|
|
42
|
+
Delegate any L0 or L1 headless-browser reproduction through
|
|
43
|
+
[UI verification](../axstack/references/ui-verification.md).
|
|
42
44
|
2. **Reproduce and minimise.** Confirm the loop reproduces the user's failure
|
|
43
45
|
and not a neighbour. Remove inputs, callers, config, data, and steps one at
|
|
44
46
|
a time within a stated budget until the repro is the smallest practical;
|
|
@@ -5,11 +5,14 @@ rendering matters.
|
|
|
5
5
|
|
|
6
6
|
1. Identify the final artifact bytes and theme. The explicit user theme wins;
|
|
7
7
|
otherwise use the dark default.
|
|
8
|
-
2.
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
8
|
+
2. Delegate the rendered pass through [UI verification](../../axstack/references/ui-verification.md).
|
|
9
|
+
Record actual desktop and mobile observations, or name the missing layout
|
|
10
|
+
check.
|
|
11
|
+
3. Have the verifier exercise relevant interactions, keyboard and screen-reader
|
|
12
|
+
accessibility, and reduced-motion behavior. Report each unavailable check
|
|
13
|
+
honestly.
|
|
14
|
+
4. The explainer reviewer checks text and source fidelity. Keep source
|
|
15
|
+
correctness, tests, rendered behavior, independent review, and publication
|
|
16
|
+
evidence separate.
|
|
14
17
|
5. Bind review to the exact artifact identity. Any byte change invalidates the
|
|
15
18
|
affected approval and requires fresh QA and review.
|
|
@@ -148,8 +148,9 @@ including evidence and retained complexity.
|
|
|
148
148
|
|
|
149
149
|
Run the acceptance checks and affected integration boundaries. Record commands,
|
|
150
150
|
observed outputs, and verified states. UI work includes rendered interaction
|
|
151
|
-
evidence
|
|
152
|
-
|
|
151
|
+
evidence through [UI verification](../axstack/references/ui-verification.md)
|
|
152
|
+
when relevant. Name every unavailable OS, harness, credential, or other
|
|
153
|
+
boundary instead of implying coverage.
|
|
153
154
|
|
|
154
155
|
After the last change, pin the exact candidate revision and return this compact
|
|
155
156
|
implementation receipt to the driver:
|
|
@@ -209,7 +209,7 @@ This section applies to peer and authored PR modes.
|
|
|
209
209
|
| `mixed` | Codex / Sol (`codex/gpt-6-sol`) | `axstack-reviewer-secondary` (`claude/claude-opus-5-5` medium) |
|
|
210
210
|
| `mixed` | Claude / Opus (`claude/claude-opus-5-5`) | `axstack-reviewer-primary` (`codex/gpt-6-sol` high) |
|
|
211
211
|
| `codex-only` | Codex / Sol (`codex/gpt-6-sol`) | `axstack-reviewer-secondary` (`codex/gpt-6-luna` xhigh) |
|
|
212
|
-
| `claude-only` | Claude / Opus (`claude/claude-opus-5-5`) | `axstack-reviewer-secondary` (`claude/claude-sonnet-5-5`
|
|
212
|
+
| `claude-only` | Claude / Opus (`claude/claude-opus-5-5`) | `axstack-reviewer-secondary` (`claude/claude-sonnet-5-5` high) |
|
|
213
213
|
|
|
214
214
|
Provenance is matched on provider/model ID; record effort, but never use
|
|
215
215
|
effort to create a mapping. Any other author provenance for the
|
|
@@ -222,7 +222,7 @@ This section applies to peer and authored PR modes.
|
|
|
222
222
|
Mixed preset review is cross-provider. Single-provider review uses the
|
|
223
223
|
configured different models and is not cross-provider independence. The
|
|
224
224
|
claude-only Sonnet explanation author/reviewer exception is session
|
|
225
|
-
independence only: separate `axstack-explainer` at
|
|
225
|
+
independence only: separate `axstack-explainer` at high and
|
|
226
226
|
`axstack-explainer-review` at high. It never permits same-model code review.
|
|
227
227
|
|
|
228
228
|
For the existing high-stakes Opus high author / Sol high checkpoint route,
|
|
@@ -279,7 +279,8 @@ This section applies to peer and authored PR modes.
|
|
|
279
279
|
|
|
280
280
|
Verify the applicable spec, ticket, or intent acceptance, executable
|
|
281
281
|
evidence, exact candidate SHA, current base, and affected integration
|
|
282
|
-
boundary, plus rendered interaction evidence for relevant UI work
|
|
282
|
+
boundary, plus rendered interaction evidence for relevant UI work through
|
|
283
|
+
[UI verification](../axstack/references/ui-verification.md). A
|
|
283
284
|
passing test is insufficient when it checks the wrong behavior. Call out
|
|
284
285
|
seeded regressions, inadequate checks, and every unverified boundary. Every
|
|
285
286
|
mode-required receipt records concrete evidence and consequences, coverage,
|
package/src/roles.js
CHANGED
|
@@ -13,7 +13,7 @@ const AUTHORED_ROUTES = Object.freeze({
|
|
|
13
13
|
'codex/gpt-6-sol': ['axstack-reviewer-secondary', 'codex/gpt-6-luna', 'xhigh'],
|
|
14
14
|
},
|
|
15
15
|
'claude-only': {
|
|
16
|
-
'claude/claude-opus-5-5': ['axstack-reviewer-secondary', 'claude/claude-sonnet-5-5', '
|
|
16
|
+
'claude/claude-opus-5-5': ['axstack-reviewer-secondary', 'claude/claude-sonnet-5-5', 'high'],
|
|
17
17
|
},
|
|
18
18
|
});
|
|
19
19
|
|