axstack 0.20.28 → 0.20.29
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +4 -1
- package/docs/installation.md +6 -2
- package/docs/workflows.md +5 -1
- package/package.json +1 -1
- package/profiles/presets/claude-only.json +15 -6
- package/profiles/presets/codex-only.json +10 -1
- package/profiles/presets/mixed.json +19 -10
- package/skills/axstack/references/orca-runtime.md +1 -1
- package/skills/axstack/references/routing.md +13 -11
- package/skills/axstack/references/ui-verification.md +13 -0
- package/skills/axstack-debug/SKILL.md +2 -0
- package/skills/axstack-explain/references/visual-qa.md +9 -6
- package/skills/axstack-implement/SKILL.md +3 -2
- package/skills/axstack-review/SKILL.md +2 -1
package/README.md
CHANGED
|
@@ -100,13 +100,16 @@ upgrades, conflicts, and uninstalling.
|
|
|
100
100
|
- Agents keep accepted decisions and evidence for resume. Missing authority,
|
|
101
101
|
unavailable models, and serious risks surface as holds. The human merges by default.
|
|
102
102
|
|
|
103
|
-
Choose one explicit preset (
|
|
103
|
+
Choose one explicit preset (28 roles each): [mixed](profiles/presets/mixed.json)
|
|
104
104
|
(recommended), [codex-only](profiles/presets/codex-only.json), or
|
|
105
105
|
[claude-only](profiles/presets/claude-only.json). Mixed supports cross-provider
|
|
106
106
|
implementation review; single-provider presets have workflow limits and are
|
|
107
107
|
not automatic fallbacks when a model is unavailable. See
|
|
108
108
|
[workflow and routing details](docs/workflows.md).
|
|
109
109
|
|
|
110
|
+
In `mixed` and `claude-only`, requirements research, web research, and the
|
|
111
|
+
optional monitor use Claude Sonnet 5.5 high.
|
|
112
|
+
|
|
110
113
|
## Optional PR automation
|
|
111
114
|
|
|
112
115
|
Manual review works without a schedule. Own open PRs in chat-run mode use a
|
package/docs/installation.md
CHANGED
|
@@ -71,7 +71,7 @@ profiles/presets/codex-only.json
|
|
|
71
71
|
profiles/presets/claude-only.json
|
|
72
72
|
```
|
|
73
73
|
|
|
74
|
-
Each has exactly `{ "version": 1, "roles": [...] }` with the same
|
|
74
|
+
Each has exactly `{ "version": 1, "roles": [...] }` with the same 28 stable
|
|
75
75
|
role IDs. Installation writes `<skills-dir>/axstack/roles.json` as
|
|
76
76
|
`{ "version": 1, "preset": "<selected preset>", "roles": [...] }` and records
|
|
77
77
|
its ownership hash like every other installed skill asset. There is no second
|
|
@@ -172,10 +172,14 @@ to rewrite them.
|
|
|
172
172
|
## Role behavior after installation
|
|
173
173
|
|
|
174
174
|
The runtime reads `roles.json` from the installed shared root `skills/axstack/`.
|
|
175
|
-
A new run records the selected preset plus all
|
|
175
|
+
A new run records the selected preset plus all 28 role rows. An active run keeps
|
|
176
176
|
that snapshot after a later preset install unless the user explicitly changes
|
|
177
177
|
it and accepts the resulting evidence invalidation.
|
|
178
178
|
|
|
179
|
+
The `mixed` and `claude-only` presets assign `axstack-research-requirements`,
|
|
180
|
+
`axstack-research-web`, and `axstack-monitor` to Claude Sonnet 5.5 high.
|
|
181
|
+
The `codex-only` assignments for these roles are unchanged.
|
|
182
|
+
|
|
179
183
|
The mixed checker and `axstack-research-web-google` have provider
|
|
180
184
|
`antigravity`; mixed `axstack-research-x` has provider `grok`. All three use
|
|
181
185
|
`model: null` because Orca exposes no model override for those agent-ID routes;
|
package/docs/workflows.md
CHANGED
|
@@ -48,7 +48,7 @@ only affected work.
|
|
|
48
48
|
|
|
49
49
|
Installation requires one explicit canonical preset. The three bundle files
|
|
50
50
|
under `profiles/presets/` each contain exactly
|
|
51
|
-
`{ "version": 1, "roles": [...] }` and the same
|
|
51
|
+
`{ "version": 1, "roles": [...] }` and the same 28 stable IDs.
|
|
52
52
|
|
|
53
53
|
The current chat drives on whatever model runs it; no preset carries a driver
|
|
54
54
|
role.
|
|
@@ -59,6 +59,10 @@ role.
|
|
|
59
59
|
| `codex-only` | Sol high | Sol high; Luna xhigh | Astra high / unavailable | Luna xhigh |
|
|
60
60
|
| `claude-only` | Opus medium | Opus medium; Sonnet high | unavailable / Opus xhigh | Sonnet high |
|
|
61
61
|
|
|
62
|
+
In `mixed` and `claude-only`, `axstack-research-requirements`,
|
|
63
|
+
`axstack-research-web`, and `axstack-monitor` use Claude Sonnet 5.5 high.
|
|
64
|
+
`codex-only` keeps its Codex assignments for those roles.
|
|
65
|
+
|
|
62
66
|
The installed `<skills-dir>/axstack/roles.json` adds the selected preset name:
|
|
63
67
|
`{ "version": 1, "preset": "<name>", "roles": [...] }`. The runtime reads it
|
|
64
68
|
from the installed shared root `skills/axstack/` and records the whole table for
|
package/package.json
CHANGED
|
@@ -68,10 +68,10 @@
|
|
|
68
68
|
"id": "axstack-research-requirements",
|
|
69
69
|
"name": "Axstack research requirements",
|
|
70
70
|
"provider": "claude",
|
|
71
|
-
"model": "claude-
|
|
71
|
+
"model": "claude-sonnet-5-5",
|
|
72
72
|
"modeId": "bypassPermissions",
|
|
73
|
-
"thinkingOptionId": "
|
|
74
|
-
"notes": "
|
|
73
|
+
"thinkingOptionId": "high",
|
|
74
|
+
"notes": "Sonnet high research requirements analyst: scopes bounded questions and acceptance for a research task. Validate configured availability at launch; hold affected work without fallback."
|
|
75
75
|
},
|
|
76
76
|
{
|
|
77
77
|
"id": "axstack-research-code",
|
|
@@ -89,7 +89,7 @@
|
|
|
89
89
|
"model": "claude-sonnet-5-5",
|
|
90
90
|
"modeId": "bypassPermissions",
|
|
91
91
|
"thinkingOptionId": "high",
|
|
92
|
-
"notes": "
|
|
92
|
+
"notes": "Sonnet high research web reader: gathers primary-source facts efficiently. Validate configured availability at launch; hold affected work without fallback."
|
|
93
93
|
},
|
|
94
94
|
{
|
|
95
95
|
"id": "axstack-research-web-google",
|
|
@@ -125,7 +125,16 @@
|
|
|
125
125
|
"model": "claude-sonnet-5-5",
|
|
126
126
|
"modeId": "bypassPermissions",
|
|
127
127
|
"thinkingOptionId": "high",
|
|
128
|
-
"notes": "Independent visual explanation reviewer: checks the exact artifact for source fidelity
|
|
128
|
+
"notes": "Independent visual explanation reviewer: checks the exact artifact for text and source fidelity. The rendered pass belongs to axstack-ui-verifier. Any artifact change invalidates its review. Validate configured availability at launch; hold affected work without fallback."
|
|
129
|
+
},
|
|
130
|
+
{
|
|
131
|
+
"id": "axstack-ui-verifier",
|
|
132
|
+
"name": "Axstack UI verifier",
|
|
133
|
+
"provider": "claude",
|
|
134
|
+
"model": "claude-sonnet-5-5",
|
|
135
|
+
"modeId": "bypassPermissions",
|
|
136
|
+
"thinkingOptionId": "high",
|
|
137
|
+
"notes": "Read-only UI verifier: runs Playwright or the browser against the given build, URL, or artifact; captures screenshots, interactions, accessibility, desktop/mobile, and reduced-motion evidence in the dispatch evidence folder; returns a verdict with evidence paths. Never edits source. Validate configured availability at launch; hold affected work without fallback."
|
|
129
138
|
},
|
|
130
139
|
{
|
|
131
140
|
"id": "axstack-explore-codebase",
|
|
@@ -152,7 +161,7 @@
|
|
|
152
161
|
"model": "claude-sonnet-5-5",
|
|
153
162
|
"modeId": "bypassPermissions",
|
|
154
163
|
"thinkingOptionId": "high",
|
|
155
|
-
"notes": "Optional independent read-only observer for a standalone PR watch. Reads GitHub, feedback, and checks, persists event IDs, and wakes the owner only for a new actionable event. Never sends, authors, reviews, replies, or acts as either reusable PR manager. Healthy snapshots stay quiet. Chat-run mode: one same-host native read-only observer per Run; fresh finite passes report precise deltas internally to the original Run/driver and may disable/read back only their own automation at verified stop. No repair, dispatch, public notification, or replacement coordinator. Effective scheduled model/effort and wake require live proof."
|
|
164
|
+
"notes": "Optional Sonnet high independent read-only observer for a standalone PR watch. Reads GitHub, feedback, and checks, persists event IDs, and wakes the owner only for a new actionable event. Never sends, authors, reviews, replies, or acts as either reusable PR manager. Healthy snapshots stay quiet. Chat-run mode: one same-host native read-only observer per Run; fresh finite passes report precise deltas internally to the original Run/driver and may disable/read back only their own automation at verified stop. No repair, dispatch, public notification, or replacement coordinator. Effective scheduled model/effort and wake require live proof."
|
|
156
165
|
},
|
|
157
166
|
{
|
|
158
167
|
"id": "axstack-auditor",
|
|
@@ -125,7 +125,16 @@
|
|
|
125
125
|
"model": "gpt-6-luna",
|
|
126
126
|
"modeId": "full-access",
|
|
127
127
|
"thinkingOptionId": "xhigh",
|
|
128
|
-
"notes": "Independent visual explanation reviewer: checks the exact artifact for source fidelity
|
|
128
|
+
"notes": "Independent visual explanation reviewer: checks the exact artifact for text and source fidelity. The rendered pass belongs to axstack-ui-verifier. Any artifact change invalidates its review. Validate configured availability at launch; hold affected work without fallback."
|
|
129
|
+
},
|
|
130
|
+
{
|
|
131
|
+
"id": "axstack-ui-verifier",
|
|
132
|
+
"name": "Axstack UI verifier",
|
|
133
|
+
"provider": "codex",
|
|
134
|
+
"model": "gpt-6-sol",
|
|
135
|
+
"modeId": "full-access",
|
|
136
|
+
"thinkingOptionId": "medium",
|
|
137
|
+
"notes": "Read-only UI verifier: runs Playwright or the browser against the given build, URL, or artifact; captures screenshots, interactions, accessibility, desktop/mobile, and reduced-motion evidence in the dispatch evidence folder; returns a verdict with evidence paths. Never edits source. Validate configured availability at launch; hold affected work without fallback."
|
|
129
138
|
},
|
|
130
139
|
{
|
|
131
140
|
"id": "axstack-explore-codebase",
|
|
@@ -68,10 +68,10 @@
|
|
|
68
68
|
"id": "axstack-research-requirements",
|
|
69
69
|
"name": "Axstack research requirements",
|
|
70
70
|
"provider": "claude",
|
|
71
|
-
"model": "claude-
|
|
71
|
+
"model": "claude-sonnet-5-5",
|
|
72
72
|
"modeId": "bypassPermissions",
|
|
73
|
-
"thinkingOptionId": "
|
|
74
|
-
"notes": "
|
|
73
|
+
"thinkingOptionId": "high",
|
|
74
|
+
"notes": "Sonnet high research requirements analyst: scopes bounded questions and acceptance for a research task. Validate configured availability at launch; hold affected work without fallback."
|
|
75
75
|
},
|
|
76
76
|
{
|
|
77
77
|
"id": "axstack-research-code",
|
|
@@ -86,10 +86,10 @@
|
|
|
86
86
|
"id": "axstack-research-web",
|
|
87
87
|
"name": "Axstack research web",
|
|
88
88
|
"provider": "claude",
|
|
89
|
-
"model": "claude-
|
|
89
|
+
"model": "claude-sonnet-5-5",
|
|
90
90
|
"modeId": "bypassPermissions",
|
|
91
|
-
"thinkingOptionId": "
|
|
92
|
-
"notes": "
|
|
91
|
+
"thinkingOptionId": "high",
|
|
92
|
+
"notes": "Sonnet high research web reader: gathers primary-source facts efficiently. Validate configured availability at launch; hold affected work without fallback."
|
|
93
93
|
},
|
|
94
94
|
{
|
|
95
95
|
"id": "axstack-research-web-google",
|
|
@@ -125,7 +125,16 @@
|
|
|
125
125
|
"model": "gpt-6-luna",
|
|
126
126
|
"modeId": "full-access",
|
|
127
127
|
"thinkingOptionId": "xhigh",
|
|
128
|
-
"notes": "Independent visual explanation reviewer: checks the exact artifact for source fidelity
|
|
128
|
+
"notes": "Independent visual explanation reviewer: checks the exact artifact for text and source fidelity. The rendered pass belongs to axstack-ui-verifier. Any artifact change invalidates its review. Validate configured availability at launch; hold affected work without fallback."
|
|
129
|
+
},
|
|
130
|
+
{
|
|
131
|
+
"id": "axstack-ui-verifier",
|
|
132
|
+
"name": "Axstack UI verifier",
|
|
133
|
+
"provider": "claude",
|
|
134
|
+
"model": "claude-sonnet-5-5",
|
|
135
|
+
"modeId": "bypassPermissions",
|
|
136
|
+
"thinkingOptionId": "high",
|
|
137
|
+
"notes": "Read-only UI verifier: runs Playwright or the browser against the given build, URL, or artifact; captures screenshots, interactions, accessibility, desktop/mobile, and reduced-motion evidence in the dispatch evidence folder; returns a verdict with evidence paths. Never edits source. Validate configured availability at launch; hold affected work without fallback."
|
|
129
138
|
},
|
|
130
139
|
{
|
|
131
140
|
"id": "axstack-explore-codebase",
|
|
@@ -149,10 +158,10 @@
|
|
|
149
158
|
"id": "axstack-monitor",
|
|
150
159
|
"name": "Axstack monitor",
|
|
151
160
|
"provider": "claude",
|
|
152
|
-
"model": "claude-
|
|
161
|
+
"model": "claude-sonnet-5-5",
|
|
153
162
|
"modeId": "bypassPermissions",
|
|
154
|
-
"thinkingOptionId": "
|
|
155
|
-
"notes": "Optional independent read-only observer for a standalone PR watch. Reads GitHub, feedback, and checks, persists event IDs, and wakes the owner only for a new actionable event. Never sends, authors, reviews, replies, or acts as either reusable PR manager. Healthy snapshots stay quiet. Chat-run mode: one same-host native read-only observer per Run; fresh finite passes report precise deltas internally to the original Run/driver and may disable/read back only their own automation at verified stop. No repair, dispatch, public notification, or replacement coordinator. Effective scheduled model/effort and wake require live proof."
|
|
163
|
+
"thinkingOptionId": "high",
|
|
164
|
+
"notes": "Optional Sonnet high independent read-only observer for a standalone PR watch. Reads GitHub, feedback, and checks, persists event IDs, and wakes the owner only for a new actionable event. Never sends, authors, reviews, replies, or acts as either reusable PR manager. Healthy snapshots stay quiet. Chat-run mode: one same-host native read-only observer per Run; fresh finite passes report precise deltas internally to the original Run/driver and may disable/read back only their own automation at verified stop. No repair, dispatch, public notification, or replacement coordinator. Effective scheduled model/effort and wake require live proof."
|
|
156
165
|
},
|
|
157
166
|
{
|
|
158
167
|
"id": "axstack-auditor",
|
|
@@ -38,7 +38,7 @@ Read `roles.json` from the installed shared root `skills/axstack/`. The installe
|
|
|
38
38
|
shape is `{ "version": 1, "preset": "<name>", "roles": [...] }`. Bundled
|
|
39
39
|
profiles are setup inputs shaped as
|
|
40
40
|
`{ "version": 1, "roles": [...] }`. A new run records the selected preset and
|
|
41
|
-
all
|
|
41
|
+
all 28 role rows once. An active run keeps the exact snapshot until the user
|
|
42
42
|
explicitly changes it.
|
|
43
43
|
|
|
44
44
|
Select the requested role by stable ID. A missing or null model holds only that role;
|
|
@@ -11,10 +11,10 @@ skills root, or an explicit user selection in the run record. Missing or contrad
|
|
|
11
11
|
a setup gap: hold. Never infer from live profiles or `list_profiles`, harness,
|
|
12
12
|
tools, credentials, quota, subscription, or default to `mixed`.
|
|
13
13
|
|
|
14
|
-
At start, snapshot all
|
|
14
|
+
At start, snapshot all 28 role IDs with provider/model/mode/effort; absent
|
|
15
15
|
or unconfigured roles are recorded explicitly; never default.
|
|
16
|
-
Such a role holds only
|
|
17
|
-
|
|
16
|
+
Such a role holds only its work. Later installed or changed roles need an
|
|
17
|
+
explicit user decision to enter the snapshot. Live profiles
|
|
18
18
|
are authoritative at snapshot time and for availability; bundled presets are setup
|
|
19
19
|
inputs, not runtime proof.
|
|
20
20
|
|
|
@@ -42,6 +42,9 @@ provider/model/effort substitution.
|
|
|
42
42
|
High-stakes/trigger: fresh [contract](contracts.md) session.
|
|
43
43
|
`axstack-auditor` audits; `axstack-checker` reports discrepancies.
|
|
44
44
|
- `axstack-explainer`/`axstack-explainer-review`: explain/review.
|
|
45
|
+
- `axstack-ui-verifier`: [UI checks](ui-verification.md).
|
|
46
|
+
- `axstack-research-requirements`/`axstack-research-web`/`axstack-monitor`:
|
|
47
|
+
Sonnet 5.5 high in mixed/claude-only.
|
|
45
48
|
`axstack-monitor`: standalone watch never sends; chat-run watch: bounded
|
|
46
49
|
internal reports to its Run and original driver.
|
|
47
50
|
- `axstack-debug-investigator-1..4` probe L1 briefs.
|
|
@@ -55,18 +58,17 @@ step (3) for user routing: no substitution or same-provider review.
|
|
|
55
58
|
|
|
56
59
|
## Direct routes (no spec ceremony)
|
|
57
60
|
|
|
58
|
-
-
|
|
59
|
-
|
|
60
|
-
-
|
|
61
|
-
current/intended behavior
|
|
62
|
-
|
|
61
|
+
- Bounded research -> `axstack-research`: verify primary sources and code,
|
|
62
|
+
cite limits, and fan out distinct questions.
|
|
63
|
+
- Understand a system or gap -> `axstack-explain`:
|
|
64
|
+
current/intended behavior and bounded gaps from docs and renders.
|
|
65
|
+
"What could this break" follows
|
|
63
66
|
[Blast radius](blast-radius.md). Publication needs separate authority.
|
|
64
67
|
- A bug, failing test, regression, or wrong behavior, red loop wanted ->
|
|
65
68
|
`axstack-debug`: diagnose, escalate via adviser-directed investigators, hand
|
|
66
69
|
off a classified repair (explain: how; debug: what's wrong).
|
|
67
|
-
-
|
|
68
|
-
|
|
69
|
-
edits.
|
|
70
|
+
- Code quality/refactor discovery -> `axstack-improve`: rank bounded
|
|
71
|
+
candidates with evidence; report only, no source edits.
|
|
70
72
|
- Accepted worker/Task/Run completion or bounded backlog request -> driver invokes
|
|
71
73
|
`axstack-cleanup` inline; never dispatch it.
|
|
72
74
|
- Preparation completion, watch expiry, resume, or reconciliation -> the
|
|
@@ -0,0 +1,13 @@
|
|
|
1
|
+
# UI verification
|
|
2
|
+
|
|
3
|
+
Every Playwright, browser, or rendered-UI check, including a "confirm it in the
|
|
4
|
+
browser" step, goes through an Orca Dispatch to `axstack-ui-verifier` from the
|
|
5
|
+
run's role snapshot. Give it the exact build, URL, or artifact and the private
|
|
6
|
+
dispatch's evidence folder. The verifier is read-only: it never edits source.
|
|
7
|
+
The PR writer remains the sole writer.
|
|
8
|
+
|
|
9
|
+
Ask for screenshots and observed interactions, accessibility, desktop and
|
|
10
|
+
mobile layouts, and reduced-motion behavior where relevant. The verifier
|
|
11
|
+
returns a verdict with evidence paths and names checks it could not run.
|
|
12
|
+
Keep the verdict tied to the exact artifact or revision; changed bytes need a
|
|
13
|
+
fresh rendered pass.
|
|
@@ -39,6 +39,8 @@ recorded reason.
|
|
|
39
39
|
loop cannot be built, stop, list what was tried, and ask the user for an
|
|
40
40
|
environment, a redacted artifact, or instrumentation permission. Done when
|
|
41
41
|
the command has run once and its red output is recorded.
|
|
42
|
+
Delegate any L0 or L1 headless-browser reproduction through
|
|
43
|
+
[UI verification](../axstack/references/ui-verification.md).
|
|
42
44
|
2. **Reproduce and minimise.** Confirm the loop reproduces the user's failure
|
|
43
45
|
and not a neighbour. Remove inputs, callers, config, data, and steps one at
|
|
44
46
|
a time within a stated budget until the repro is the smallest practical;
|
|
@@ -5,11 +5,14 @@ rendering matters.
|
|
|
5
5
|
|
|
6
6
|
1. Identify the final artifact bytes and theme. The explicit user theme wins;
|
|
7
7
|
otherwise use the dark default.
|
|
8
|
-
2.
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
8
|
+
2. Delegate the rendered pass through [UI verification](../../axstack/references/ui-verification.md).
|
|
9
|
+
Record actual desktop and mobile observations, or name the missing layout
|
|
10
|
+
check.
|
|
11
|
+
3. Have the verifier exercise relevant interactions, keyboard and screen-reader
|
|
12
|
+
accessibility, and reduced-motion behavior. Report each unavailable check
|
|
13
|
+
honestly.
|
|
14
|
+
4. The explainer reviewer checks text and source fidelity. Keep source
|
|
15
|
+
correctness, tests, rendered behavior, independent review, and publication
|
|
16
|
+
evidence separate.
|
|
14
17
|
5. Bind review to the exact artifact identity. Any byte change invalidates the
|
|
15
18
|
affected approval and requires fresh QA and review.
|
|
@@ -148,8 +148,9 @@ including evidence and retained complexity.
|
|
|
148
148
|
|
|
149
149
|
Run the acceptance checks and affected integration boundaries. Record commands,
|
|
150
150
|
observed outputs, and verified states. UI work includes rendered interaction
|
|
151
|
-
evidence
|
|
152
|
-
|
|
151
|
+
evidence through [UI verification](../axstack/references/ui-verification.md)
|
|
152
|
+
when relevant. Name every unavailable OS, harness, credential, or other
|
|
153
|
+
boundary instead of implying coverage.
|
|
153
154
|
|
|
154
155
|
After the last change, pin the exact candidate revision and return this compact
|
|
155
156
|
implementation receipt to the driver:
|
|
@@ -279,7 +279,8 @@ This section applies to peer and authored PR modes.
|
|
|
279
279
|
|
|
280
280
|
Verify the applicable spec, ticket, or intent acceptance, executable
|
|
281
281
|
evidence, exact candidate SHA, current base, and affected integration
|
|
282
|
-
boundary, plus rendered interaction evidence for relevant UI work
|
|
282
|
+
boundary, plus rendered interaction evidence for relevant UI work through
|
|
283
|
+
[UI verification](../axstack/references/ui-verification.md). A
|
|
283
284
|
passing test is insufficient when it checks the wrong behavior. Call out
|
|
284
285
|
seeded regressions, inadequate checks, and every unverified boundary. Every
|
|
285
286
|
mode-required receipt records concrete evidence and consequences, coverage,
|