axstack 0.20.28 → 0.20.30
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +6 -1
- package/docs/installation.md +11 -2
- package/docs/workflows.md +14 -4
- package/package.json +1 -1
- package/profiles/presets/claude-only.json +56 -11
- package/profiles/presets/codex-only.json +46 -1
- package/profiles/presets/mixed.json +60 -15
- package/skills/axstack/references/candidate-publication.md +9 -0
- package/skills/axstack/references/diligence.md +23 -0
- package/skills/axstack/references/lifecycle.md +2 -2
- package/skills/axstack/references/orca-runtime.md +1 -1
- package/skills/axstack/references/role-roster.md +37 -0
- package/skills/axstack/references/routing.md +11 -32
- package/skills/axstack/references/ui-verification.md +13 -0
- package/skills/axstack-audit/SKILL.md +7 -2
- package/skills/axstack-debug/SKILL.md +2 -0
- package/skills/axstack-explain/references/visual-qa.md +9 -6
- package/skills/axstack-implement/SKILL.md +10 -4
- package/skills/axstack-improve/SKILL.md +5 -0
- package/skills/axstack-research/SKILL.md +14 -2
- package/skills/axstack-review/SKILL.md +18 -7
- package/skills/axstack-spec/SKILL.md +3 -0
- package/skills/axstack-tickets/SKILL.md +3 -0
- package/skills/axstack-watch/SKILL.md +1 -0
- package/src/roles.js +1 -1
package/README.md
CHANGED
|
@@ -100,13 +100,18 @@ upgrades, conflicts, and uninstalling.
|
|
|
100
100
|
- Agents keep accepted decisions and evidence for resume. Missing authority,
|
|
101
101
|
unavailable models, and serious risks surface as holds. The human merges by default.
|
|
102
102
|
|
|
103
|
-
Choose one explicit preset (
|
|
103
|
+
Choose one explicit preset (32 roles each): [mixed](profiles/presets/mixed.json)
|
|
104
104
|
(recommended), [codex-only](profiles/presets/codex-only.json), or
|
|
105
105
|
[claude-only](profiles/presets/claude-only.json). Mixed supports cross-provider
|
|
106
106
|
implementation review; single-provider presets have workflow limits and are
|
|
107
107
|
not automatic fallbacks when a model is unavailable. See
|
|
108
108
|
[workflow and routing details](docs/workflows.md).
|
|
109
109
|
|
|
110
|
+
In `mixed` and `claude-only`, auditing, requirements/code/web research,
|
|
111
|
+
execution exploration, and the optional monitor use Claude Sonnet 5.5 high.
|
|
112
|
+
Auditing, code research, and execution exploration have independent Sol high
|
|
113
|
+
pair seats in `mixed` and `codex-only`; `claude-only` records them as absent.
|
|
114
|
+
|
|
110
115
|
## Optional PR automation
|
|
111
116
|
|
|
112
117
|
Manual review works without a schedule. Own open PRs in chat-run mode use a
|
package/docs/installation.md
CHANGED
|
@@ -71,7 +71,7 @@ profiles/presets/codex-only.json
|
|
|
71
71
|
profiles/presets/claude-only.json
|
|
72
72
|
```
|
|
73
73
|
|
|
74
|
-
Each has exactly `{ "version": 1, "roles": [...] }` with the same
|
|
74
|
+
Each has exactly `{ "version": 1, "roles": [...] }` with the same 32 stable
|
|
75
75
|
role IDs. Installation writes `<skills-dir>/axstack/roles.json` as
|
|
76
76
|
`{ "version": 1, "preset": "<selected preset>", "roles": [...] }` and records
|
|
77
77
|
its ownership hash like every other installed skill asset. There is no second
|
|
@@ -172,10 +172,19 @@ to rewrite them.
|
|
|
172
172
|
## Role behavior after installation
|
|
173
173
|
|
|
174
174
|
The runtime reads `roles.json` from the installed shared root `skills/axstack/`.
|
|
175
|
-
A new run records the selected preset plus all
|
|
175
|
+
A new run records the selected preset plus all 32 role rows. An active run keeps
|
|
176
176
|
that snapshot after a later preset install unless the user explicitly changes
|
|
177
177
|
it and accepts the resulting evidence invalidation.
|
|
178
178
|
|
|
179
|
+
The `mixed` and `claude-only` presets assign `axstack-auditor`,
|
|
180
|
+
`axstack-research-requirements`, `axstack-research-code`, `axstack-research-web`,
|
|
181
|
+
`axstack-explore-execution`, and `axstack-monitor` to Claude Sonnet 5.5 high.
|
|
182
|
+
The `codex-only` assignments for these roles are unchanged.
|
|
183
|
+
The three `-sol` pair seats for auditor, research-code, and explore-execution
|
|
184
|
+
use Sol high in `mixed` and `codex-only`; `claude-only` records intentional
|
|
185
|
+
absences. The paired seats run independently on one brief and the driver
|
|
186
|
+
reconciles their findings.
|
|
187
|
+
|
|
179
188
|
The mixed checker and `axstack-research-web-google` have provider
|
|
180
189
|
`antigravity`; mixed `axstack-research-x` has provider `grok`. All three use
|
|
181
190
|
`model: null` because Orca exposes no model override for those agent-ID routes;
|
package/docs/workflows.md
CHANGED
|
@@ -48,16 +48,26 @@ only affected work.
|
|
|
48
48
|
|
|
49
49
|
Installation requires one explicit canonical preset. The three bundle files
|
|
50
50
|
under `profiles/presets/` each contain exactly
|
|
51
|
-
`{ "version": 1, "roles": [...] }` and the same
|
|
51
|
+
`{ "version": 1, "roles": [...] }` and the same 32 stable IDs.
|
|
52
52
|
|
|
53
53
|
The current chat drives on whatever model runs it; no preset carries a driver
|
|
54
54
|
role.
|
|
55
55
|
|
|
56
56
|
| Preset | Author | Ordered peer reviewers | Astra / Opus advisers | Auditor |
|
|
57
57
|
| --- | --- | --- | --- | --- |
|
|
58
|
-
| `mixed` | Sol high | Sol high; Opus medium | Astra high / Opus xhigh |
|
|
59
|
-
| `codex-only` | Sol high | Sol high; Luna xhigh | Astra high / unavailable | Luna xhigh |
|
|
60
|
-
| `claude-only` | Opus medium | Opus medium; Sonnet high | unavailable / Opus xhigh | Sonnet high |
|
|
58
|
+
| `mixed` | Sol high | Sol high; Opus medium | Astra high / Opus xhigh | Sonnet high + Sol high |
|
|
59
|
+
| `codex-only` | Sol high | Sol high; Luna xhigh | Astra high / unavailable | Luna xhigh + Sol high |
|
|
60
|
+
| `claude-only` | Opus medium | Opus medium; Sonnet high | unavailable / Opus xhigh | Sonnet high (Sol absent) |
|
|
61
|
+
|
|
62
|
+
In `mixed` and `claude-only`, `axstack-auditor`, `axstack-research-requirements`,
|
|
63
|
+
`axstack-research-code`, `axstack-research-web`, `axstack-explore-execution`,
|
|
64
|
+
and `axstack-monitor` use Claude Sonnet 5.5 high. `codex-only` keeps its Codex
|
|
65
|
+
assignments for those roles. Mixed web-google and X retain their source-specific
|
|
66
|
+
Antigravity and Grok routes.
|
|
67
|
+
The new `-sol` auditor, research-code, and explore-execution seats use Sol high
|
|
68
|
+
in `mixed` and `codex-only`; `claude-only` records each as intentionally absent.
|
|
69
|
+
Dispatch the base and Sol seats independently on the same brief, then reconcile
|
|
70
|
+
their findings per claim without averaging.
|
|
61
71
|
|
|
62
72
|
The installed `<skills-dir>/axstack/roles.json` adds the selected preset name:
|
|
63
73
|
`{ "version": 1, "preset": "<name>", "roles": [...] }`. The runtime reads it
|
package/package.json
CHANGED
|
@@ -55,6 +55,15 @@
|
|
|
55
55
|
"thinkingOptionId": "high",
|
|
56
56
|
"notes": "Secondary reviewer in the ordered claude-only peer pair: Opus medium followed by Sonnet high. Eligible authored reviewer for an Opus-authored candidate. Never author or owner; review exact SHA and base across all six angles and acceptance. Peer first pass stays isolated."
|
|
57
57
|
},
|
|
58
|
+
{
|
|
59
|
+
"id": "axstack-diligence",
|
|
60
|
+
"name": "Axstack diligence checker",
|
|
61
|
+
"provider": "claude",
|
|
62
|
+
"model": "claude-sonnet-5-5",
|
|
63
|
+
"modeId": "bypassPermissions",
|
|
64
|
+
"thinkingOptionId": "high",
|
|
65
|
+
"notes": "Read-only diligence for exact-revision PRs and bounded research, spec, ticket, receipt, and release claims. Returns PASS or FINDINGS with evidence; never authors or edits."
|
|
66
|
+
},
|
|
58
67
|
{
|
|
59
68
|
"id": "axstack-checker",
|
|
60
69
|
"name": "Axstack tracker checker",
|
|
@@ -68,19 +77,28 @@
|
|
|
68
77
|
"id": "axstack-research-requirements",
|
|
69
78
|
"name": "Axstack research requirements",
|
|
70
79
|
"provider": "claude",
|
|
71
|
-
"model": "claude-
|
|
80
|
+
"model": "claude-sonnet-5-5",
|
|
72
81
|
"modeId": "bypassPermissions",
|
|
73
|
-
"thinkingOptionId": "
|
|
74
|
-
"notes": "
|
|
82
|
+
"thinkingOptionId": "high",
|
|
83
|
+
"notes": "Sonnet high research requirements analyst: scopes bounded questions and acceptance for a research task. Validate configured availability at launch; hold affected work without fallback."
|
|
75
84
|
},
|
|
76
85
|
{
|
|
77
86
|
"id": "axstack-research-code",
|
|
78
87
|
"name": "Axstack research code",
|
|
79
88
|
"provider": "claude",
|
|
80
|
-
"model": "claude-
|
|
89
|
+
"model": "claude-sonnet-5-5",
|
|
81
90
|
"modeId": "bypassPermissions",
|
|
82
|
-
"thinkingOptionId": "
|
|
83
|
-
"notes": "
|
|
91
|
+
"thinkingOptionId": "high",
|
|
92
|
+
"notes": "Sonnet high research code investigator: verifies behavior against inspected code and executable evidence. Validate configured availability at launch; hold affected work without fallback."
|
|
93
|
+
},
|
|
94
|
+
{
|
|
95
|
+
"id": "axstack-research-code-sol",
|
|
96
|
+
"name": "Axstack research code Sol (unavailable)",
|
|
97
|
+
"provider": "claude",
|
|
98
|
+
"model": null,
|
|
99
|
+
"modeId": "bypassPermissions",
|
|
100
|
+
"thinkingOptionId": "high",
|
|
101
|
+
"notes": "Intentional single-provider absence: Sol pair seat axstack-research-code-sol is unavailable in claude-only. The explicit null records the absent pair without substitution; the Sonnet seat proceeds alone."
|
|
84
102
|
},
|
|
85
103
|
{
|
|
86
104
|
"id": "axstack-research-web",
|
|
@@ -89,7 +107,7 @@
|
|
|
89
107
|
"model": "claude-sonnet-5-5",
|
|
90
108
|
"modeId": "bypassPermissions",
|
|
91
109
|
"thinkingOptionId": "high",
|
|
92
|
-
"notes": "
|
|
110
|
+
"notes": "Sonnet high research web reader: gathers primary-source facts efficiently. Validate configured availability at launch; hold affected work without fallback."
|
|
93
111
|
},
|
|
94
112
|
{
|
|
95
113
|
"id": "axstack-research-web-google",
|
|
@@ -125,7 +143,16 @@
|
|
|
125
143
|
"model": "claude-sonnet-5-5",
|
|
126
144
|
"modeId": "bypassPermissions",
|
|
127
145
|
"thinkingOptionId": "high",
|
|
128
|
-
"notes": "Independent visual explanation reviewer: checks the exact artifact for source fidelity
|
|
146
|
+
"notes": "Independent visual explanation reviewer: checks the exact artifact for text and source fidelity. The rendered pass belongs to axstack-ui-verifier. Any artifact change invalidates its review. Validate configured availability at launch; hold affected work without fallback."
|
|
147
|
+
},
|
|
148
|
+
{
|
|
149
|
+
"id": "axstack-ui-verifier",
|
|
150
|
+
"name": "Axstack UI verifier",
|
|
151
|
+
"provider": "claude",
|
|
152
|
+
"model": "claude-sonnet-5-5",
|
|
153
|
+
"modeId": "bypassPermissions",
|
|
154
|
+
"thinkingOptionId": "high",
|
|
155
|
+
"notes": "Read-only UI verifier: runs Playwright or the browser against the given build, URL, or artifact; captures screenshots, interactions, accessibility, desktop/mobile, and reduced-motion evidence in the dispatch evidence folder; returns a verdict with evidence paths. Never edits source. Validate configured availability at launch; hold affected work without fallback."
|
|
129
156
|
},
|
|
130
157
|
{
|
|
131
158
|
"id": "axstack-explore-codebase",
|
|
@@ -143,7 +170,16 @@
|
|
|
143
170
|
"model": "claude-sonnet-5-5",
|
|
144
171
|
"modeId": "bypassPermissions",
|
|
145
172
|
"thinkingOptionId": "high",
|
|
146
|
-
"notes": "
|
|
173
|
+
"notes": "Sonnet high execution explorer: runs bounded checks of runtime behavior where authorized. Validate configured availability at launch; hold affected work without fallback."
|
|
174
|
+
},
|
|
175
|
+
{
|
|
176
|
+
"id": "axstack-explore-execution-sol",
|
|
177
|
+
"name": "Axstack execution explorer Sol (unavailable)",
|
|
178
|
+
"provider": "claude",
|
|
179
|
+
"model": null,
|
|
180
|
+
"modeId": "bypassPermissions",
|
|
181
|
+
"thinkingOptionId": "high",
|
|
182
|
+
"notes": "Intentional single-provider absence: Sol pair seat axstack-explore-execution-sol is unavailable in claude-only. The explicit null records the absent pair without substitution; the Sonnet seat proceeds alone."
|
|
147
183
|
},
|
|
148
184
|
{
|
|
149
185
|
"id": "axstack-monitor",
|
|
@@ -152,7 +188,7 @@
|
|
|
152
188
|
"model": "claude-sonnet-5-5",
|
|
153
189
|
"modeId": "bypassPermissions",
|
|
154
190
|
"thinkingOptionId": "high",
|
|
155
|
-
"notes": "Optional independent read-only observer for a standalone PR watch. Reads GitHub, feedback, and checks, persists event IDs, and wakes the owner only for a new actionable event. Never sends, authors, reviews, replies, or acts as either reusable PR manager. Healthy snapshots stay quiet. Chat-run mode: one same-host native read-only observer per Run; fresh finite passes report precise deltas internally to the original Run/driver and may disable/read back only their own automation at verified stop. No repair, dispatch, public notification, or replacement coordinator. Effective scheduled model/effort and wake require live proof."
|
|
191
|
+
"notes": "Optional Sonnet high independent read-only observer for a standalone PR watch. Reads GitHub, feedback, and checks, persists event IDs, and wakes the owner only for a new actionable event. Never sends, authors, reviews, replies, or acts as either reusable PR manager. Healthy snapshots stay quiet. Chat-run mode: one same-host native read-only observer per Run; fresh finite passes report precise deltas internally to the original Run/driver and may disable/read back only their own automation at verified stop. No repair, dispatch, public notification, or replacement coordinator. Effective scheduled model/effort and wake require live proof."
|
|
156
192
|
},
|
|
157
193
|
{
|
|
158
194
|
"id": "axstack-auditor",
|
|
@@ -161,7 +197,16 @@
|
|
|
161
197
|
"model": "claude-sonnet-5-5",
|
|
162
198
|
"modeId": "bypassPermissions",
|
|
163
199
|
"thinkingOptionId": "high",
|
|
164
|
-
"notes": "
|
|
200
|
+
"notes": "Sonnet high read-only end-of-run and checkpoint auditor. Collects scope and outcome evidence with counts and denominators and reports PASS, FAIL, or UNKNOWN without inventing numbers. Never edits, merges, activates, or audits itself."
|
|
201
|
+
},
|
|
202
|
+
{
|
|
203
|
+
"id": "axstack-auditor-sol",
|
|
204
|
+
"name": "Axstack auditor Sol (unavailable)",
|
|
205
|
+
"provider": "claude",
|
|
206
|
+
"model": null,
|
|
207
|
+
"modeId": "bypassPermissions",
|
|
208
|
+
"thinkingOptionId": "high",
|
|
209
|
+
"notes": "Intentional single-provider absence: Sol pair seat axstack-auditor-sol is unavailable in claude-only. The explicit null records the absent pair without substitution; the Sonnet seat proceeds alone."
|
|
165
210
|
},
|
|
166
211
|
{
|
|
167
212
|
"id": "axstack-debug-investigator-1",
|
|
@@ -55,6 +55,15 @@
|
|
|
55
55
|
"thinkingOptionId": "xhigh",
|
|
56
56
|
"notes": "Secondary reviewer in the ordered codex-only peer pair: Sol high followed by Luna xhigh. Eligible authored reviewer for a Sol-authored candidate. Never author or owner; review exact SHA and base across all six angles and acceptance. Peer first pass stays isolated."
|
|
57
57
|
},
|
|
58
|
+
{
|
|
59
|
+
"id": "axstack-diligence",
|
|
60
|
+
"name": "Axstack diligence checker",
|
|
61
|
+
"provider": "codex",
|
|
62
|
+
"model": "gpt-6-sol",
|
|
63
|
+
"modeId": "full-access",
|
|
64
|
+
"thinkingOptionId": "high",
|
|
65
|
+
"notes": "Read-only diligence for exact-revision PRs and bounded research, spec, ticket, receipt, and release claims. Returns PASS or FINDINGS with evidence; never authors or edits."
|
|
66
|
+
},
|
|
58
67
|
{
|
|
59
68
|
"id": "axstack-checker",
|
|
60
69
|
"name": "Axstack tracker checker",
|
|
@@ -82,6 +91,15 @@
|
|
|
82
91
|
"thinkingOptionId": "high",
|
|
83
92
|
"notes": "Research code investigator: verifies behavior against inspected code and executable evidence. Validate configured availability at launch; hold affected work without fallback."
|
|
84
93
|
},
|
|
94
|
+
{
|
|
95
|
+
"id": "axstack-research-code-sol",
|
|
96
|
+
"name": "Axstack research code Sol",
|
|
97
|
+
"provider": "codex",
|
|
98
|
+
"model": "gpt-6-sol",
|
|
99
|
+
"modeId": "full-access",
|
|
100
|
+
"thinkingOptionId": "high",
|
|
101
|
+
"notes": "Independent Sol high research code investigator: verifies code behavior and executable evidence on the same bounded brief without cross-reading. Reports findings for driver reconciliation."
|
|
102
|
+
},
|
|
85
103
|
{
|
|
86
104
|
"id": "axstack-research-web",
|
|
87
105
|
"name": "Axstack research web",
|
|
@@ -125,7 +143,16 @@
|
|
|
125
143
|
"model": "gpt-6-luna",
|
|
126
144
|
"modeId": "full-access",
|
|
127
145
|
"thinkingOptionId": "xhigh",
|
|
128
|
-
"notes": "Independent visual explanation reviewer: checks the exact artifact for source fidelity
|
|
146
|
+
"notes": "Independent visual explanation reviewer: checks the exact artifact for text and source fidelity. The rendered pass belongs to axstack-ui-verifier. Any artifact change invalidates its review. Validate configured availability at launch; hold affected work without fallback."
|
|
147
|
+
},
|
|
148
|
+
{
|
|
149
|
+
"id": "axstack-ui-verifier",
|
|
150
|
+
"name": "Axstack UI verifier",
|
|
151
|
+
"provider": "codex",
|
|
152
|
+
"model": "gpt-6-sol",
|
|
153
|
+
"modeId": "full-access",
|
|
154
|
+
"thinkingOptionId": "medium",
|
|
155
|
+
"notes": "Read-only UI verifier: runs Playwright or the browser against the given build, URL, or artifact; captures screenshots, interactions, accessibility, desktop/mobile, and reduced-motion evidence in the dispatch evidence folder; returns a verdict with evidence paths. Never edits source. Validate configured availability at launch; hold affected work without fallback."
|
|
129
156
|
},
|
|
130
157
|
{
|
|
131
158
|
"id": "axstack-explore-codebase",
|
|
@@ -145,6 +172,15 @@
|
|
|
145
172
|
"thinkingOptionId": "high",
|
|
146
173
|
"notes": "Execution explorer: runs bounded checks of runtime behavior where authorized. Validate configured availability at launch; hold affected work without fallback."
|
|
147
174
|
},
|
|
175
|
+
{
|
|
176
|
+
"id": "axstack-explore-execution-sol",
|
|
177
|
+
"name": "Axstack execution explorer Sol",
|
|
178
|
+
"provider": "codex",
|
|
179
|
+
"model": "gpt-6-sol",
|
|
180
|
+
"modeId": "full-access",
|
|
181
|
+
"thinkingOptionId": "high",
|
|
182
|
+
"notes": "Independent Sol high execution explorer: runs authorized bounded runtime checks on the same brief without cross-reading. Reports findings for driver reconciliation."
|
|
183
|
+
},
|
|
148
184
|
{
|
|
149
185
|
"id": "axstack-monitor",
|
|
150
186
|
"name": "Axstack monitor",
|
|
@@ -163,6 +199,15 @@
|
|
|
163
199
|
"thinkingOptionId": "xhigh",
|
|
164
200
|
"notes": "Read-only end-of-run and checkpoint auditor. Collects scope and outcome evidence with counts and denominators and reports PASS, FAIL, or UNKNOWN without inventing numbers. Never edits, merges, activates, or audits itself."
|
|
165
201
|
},
|
|
202
|
+
{
|
|
203
|
+
"id": "axstack-auditor-sol",
|
|
204
|
+
"name": "Axstack auditor Sol",
|
|
205
|
+
"provider": "codex",
|
|
206
|
+
"model": "gpt-6-sol",
|
|
207
|
+
"modeId": "full-access",
|
|
208
|
+
"thinkingOptionId": "high",
|
|
209
|
+
"notes": "Independent Sol high read-only auditor: checks the same bounded run evidence without cross-reading. Reports findings for driver reconciliation; never edits, merges, or activates."
|
|
210
|
+
},
|
|
166
211
|
{
|
|
167
212
|
"id": "axstack-debug-investigator-1",
|
|
168
213
|
"name": "Axstack debug investigator 1",
|
|
@@ -55,6 +55,15 @@
|
|
|
55
55
|
"thinkingOptionId": "medium",
|
|
56
56
|
"notes": "Secondary reviewer in the ordered mixed peer pair: Sol high followed by Opus medium. Eligible authored reviewer for a Sol-authored candidate. Never author or owner; review exact SHA and base across all six angles and acceptance. Peer first pass stays isolated."
|
|
57
57
|
},
|
|
58
|
+
{
|
|
59
|
+
"id": "axstack-diligence",
|
|
60
|
+
"name": "Axstack diligence checker",
|
|
61
|
+
"provider": "claude",
|
|
62
|
+
"model": "claude-sonnet-5-5",
|
|
63
|
+
"modeId": "bypassPermissions",
|
|
64
|
+
"thinkingOptionId": "high",
|
|
65
|
+
"notes": "Read-only diligence for exact-revision PRs and bounded research, spec, ticket, receipt, and release claims. Returns PASS or FINDINGS with evidence; never authors or edits."
|
|
66
|
+
},
|
|
58
67
|
{
|
|
59
68
|
"id": "axstack-checker",
|
|
60
69
|
"name": "Axstack tracker checker",
|
|
@@ -68,28 +77,37 @@
|
|
|
68
77
|
"id": "axstack-research-requirements",
|
|
69
78
|
"name": "Axstack research requirements",
|
|
70
79
|
"provider": "claude",
|
|
71
|
-
"model": "claude-
|
|
80
|
+
"model": "claude-sonnet-5-5",
|
|
72
81
|
"modeId": "bypassPermissions",
|
|
73
|
-
"thinkingOptionId": "
|
|
74
|
-
"notes": "
|
|
82
|
+
"thinkingOptionId": "high",
|
|
83
|
+
"notes": "Sonnet high research requirements analyst: scopes bounded questions and acceptance for a research task. Validate configured availability at launch; hold affected work without fallback."
|
|
75
84
|
},
|
|
76
85
|
{
|
|
77
86
|
"id": "axstack-research-code",
|
|
78
87
|
"name": "Axstack research code",
|
|
88
|
+
"provider": "claude",
|
|
89
|
+
"model": "claude-sonnet-5-5",
|
|
90
|
+
"modeId": "bypassPermissions",
|
|
91
|
+
"thinkingOptionId": "high",
|
|
92
|
+
"notes": "Sonnet high research code investigator: verifies behavior against inspected code and executable evidence. Validate configured availability at launch; hold affected work without fallback."
|
|
93
|
+
},
|
|
94
|
+
{
|
|
95
|
+
"id": "axstack-research-code-sol",
|
|
96
|
+
"name": "Axstack research code Sol",
|
|
79
97
|
"provider": "codex",
|
|
80
98
|
"model": "gpt-6-sol",
|
|
81
99
|
"modeId": "full-access",
|
|
82
100
|
"thinkingOptionId": "high",
|
|
83
|
-
"notes": "
|
|
101
|
+
"notes": "Independent Sol high research code investigator: verifies code behavior and executable evidence on the same bounded brief without cross-reading. Reports findings for driver reconciliation."
|
|
84
102
|
},
|
|
85
103
|
{
|
|
86
104
|
"id": "axstack-research-web",
|
|
87
105
|
"name": "Axstack research web",
|
|
88
106
|
"provider": "claude",
|
|
89
|
-
"model": "claude-
|
|
107
|
+
"model": "claude-sonnet-5-5",
|
|
90
108
|
"modeId": "bypassPermissions",
|
|
91
|
-
"thinkingOptionId": "
|
|
92
|
-
"notes": "
|
|
109
|
+
"thinkingOptionId": "high",
|
|
110
|
+
"notes": "Sonnet high research web reader: gathers primary-source facts efficiently. Validate configured availability at launch; hold affected work without fallback."
|
|
93
111
|
},
|
|
94
112
|
{
|
|
95
113
|
"id": "axstack-research-web-google",
|
|
@@ -125,7 +143,16 @@
|
|
|
125
143
|
"model": "gpt-6-luna",
|
|
126
144
|
"modeId": "full-access",
|
|
127
145
|
"thinkingOptionId": "xhigh",
|
|
128
|
-
"notes": "Independent visual explanation reviewer: checks the exact artifact for source fidelity
|
|
146
|
+
"notes": "Independent visual explanation reviewer: checks the exact artifact for text and source fidelity. The rendered pass belongs to axstack-ui-verifier. Any artifact change invalidates its review. Validate configured availability at launch; hold affected work without fallback."
|
|
147
|
+
},
|
|
148
|
+
{
|
|
149
|
+
"id": "axstack-ui-verifier",
|
|
150
|
+
"name": "Axstack UI verifier",
|
|
151
|
+
"provider": "claude",
|
|
152
|
+
"model": "claude-sonnet-5-5",
|
|
153
|
+
"modeId": "bypassPermissions",
|
|
154
|
+
"thinkingOptionId": "high",
|
|
155
|
+
"notes": "Read-only UI verifier: runs Playwright or the browser against the given build, URL, or artifact; captures screenshots, interactions, accessibility, desktop/mobile, and reduced-motion evidence in the dispatch evidence folder; returns a verdict with evidence paths. Never edits source. Validate configured availability at launch; hold affected work without fallback."
|
|
129
156
|
},
|
|
130
157
|
{
|
|
131
158
|
"id": "axstack-explore-codebase",
|
|
@@ -139,29 +166,47 @@
|
|
|
139
166
|
{
|
|
140
167
|
"id": "axstack-explore-execution",
|
|
141
168
|
"name": "Axstack execution explorer",
|
|
169
|
+
"provider": "claude",
|
|
170
|
+
"model": "claude-sonnet-5-5",
|
|
171
|
+
"modeId": "bypassPermissions",
|
|
172
|
+
"thinkingOptionId": "high",
|
|
173
|
+
"notes": "Sonnet high execution explorer: runs bounded checks of runtime behavior where authorized. Validate configured availability at launch; hold affected work without fallback."
|
|
174
|
+
},
|
|
175
|
+
{
|
|
176
|
+
"id": "axstack-explore-execution-sol",
|
|
177
|
+
"name": "Axstack execution explorer Sol",
|
|
142
178
|
"provider": "codex",
|
|
143
179
|
"model": "gpt-6-sol",
|
|
144
180
|
"modeId": "full-access",
|
|
145
181
|
"thinkingOptionId": "high",
|
|
146
|
-
"notes": "
|
|
182
|
+
"notes": "Independent Sol high execution explorer: runs authorized bounded runtime checks on the same brief without cross-reading. Reports findings for driver reconciliation."
|
|
147
183
|
},
|
|
148
184
|
{
|
|
149
185
|
"id": "axstack-monitor",
|
|
150
186
|
"name": "Axstack monitor",
|
|
151
187
|
"provider": "claude",
|
|
152
|
-
"model": "claude-
|
|
188
|
+
"model": "claude-sonnet-5-5",
|
|
153
189
|
"modeId": "bypassPermissions",
|
|
154
|
-
"thinkingOptionId": "
|
|
155
|
-
"notes": "Optional independent read-only observer for a standalone PR watch. Reads GitHub, feedback, and checks, persists event IDs, and wakes the owner only for a new actionable event. Never sends, authors, reviews, replies, or acts as either reusable PR manager. Healthy snapshots stay quiet. Chat-run mode: one same-host native read-only observer per Run; fresh finite passes report precise deltas internally to the original Run/driver and may disable/read back only their own automation at verified stop. No repair, dispatch, public notification, or replacement coordinator. Effective scheduled model/effort and wake require live proof."
|
|
190
|
+
"thinkingOptionId": "high",
|
|
191
|
+
"notes": "Optional Sonnet high independent read-only observer for a standalone PR watch. Reads GitHub, feedback, and checks, persists event IDs, and wakes the owner only for a new actionable event. Never sends, authors, reviews, replies, or acts as either reusable PR manager. Healthy snapshots stay quiet. Chat-run mode: one same-host native read-only observer per Run; fresh finite passes report precise deltas internally to the original Run/driver and may disable/read back only their own automation at verified stop. No repair, dispatch, public notification, or replacement coordinator. Effective scheduled model/effort and wake require live proof."
|
|
156
192
|
},
|
|
157
193
|
{
|
|
158
194
|
"id": "axstack-auditor",
|
|
159
195
|
"name": "Axstack auditor",
|
|
196
|
+
"provider": "claude",
|
|
197
|
+
"model": "claude-sonnet-5-5",
|
|
198
|
+
"modeId": "bypassPermissions",
|
|
199
|
+
"thinkingOptionId": "high",
|
|
200
|
+
"notes": "Sonnet high read-only end-of-run and checkpoint auditor. Collects scope and outcome evidence with counts and denominators and reports PASS, FAIL, or UNKNOWN without inventing numbers. Never edits, merges, activates, or audits itself."
|
|
201
|
+
},
|
|
202
|
+
{
|
|
203
|
+
"id": "axstack-auditor-sol",
|
|
204
|
+
"name": "Axstack auditor Sol",
|
|
160
205
|
"provider": "codex",
|
|
161
|
-
"model": "gpt-6-
|
|
206
|
+
"model": "gpt-6-sol",
|
|
162
207
|
"modeId": "full-access",
|
|
163
|
-
"thinkingOptionId": "
|
|
164
|
-
"notes": "
|
|
208
|
+
"thinkingOptionId": "high",
|
|
209
|
+
"notes": "Independent Sol high read-only auditor: checks the same bounded run evidence without cross-reading. Reports findings for driver reconciliation; never edits, merges, or activates."
|
|
165
210
|
},
|
|
166
211
|
{
|
|
167
212
|
"id": "axstack-debug-investigator-1",
|
|
@@ -6,6 +6,11 @@ Within recorded PR-scoped publication authority, the owner reconciles that
|
|
|
6
6
|
receipt against the actual local candidate SHA and base. The owner does not edit
|
|
7
7
|
the author's candidate; required code changes return to the author.
|
|
8
8
|
|
|
9
|
+
Before publication, dispatch `axstack-diligence` under
|
|
10
|
+
[Diligence](diligence.md) to check the author receipt against its evidence
|
|
11
|
+
folder: red/green logs exist, and counts, SHAs, and paths match. Resolve
|
|
12
|
+
`FINDINGS` with the same author before publishing.
|
|
13
|
+
|
|
9
14
|
Publish the existing commits through `gh stack`. Prefer a fast-forward push.
|
|
10
15
|
Before a history rewrite, confirm the expected-old remote SHA and use lease
|
|
11
16
|
protection; a mismatch holds publication. If the push outcome is ambiguous,
|
|
@@ -40,3 +45,7 @@ as a `git clone` into a temp directory followed by `orca repo add`; each
|
|
|
40
45
|
the directory is gone. Release preparation uses a `release/<version>` worktree
|
|
41
46
|
of the same registered repo the same way. Release the checkout with
|
|
42
47
|
`ORCA worktree rm` after its receipt is recorded.
|
|
48
|
+
|
|
49
|
+
For a release PR, dispatch `axstack-diligence` under
|
|
50
|
+
[Diligence](diligence.md) to check the release PR body
|
|
51
|
+
against the merged PRs before publication.
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
# Diligence
|
|
2
|
+
|
|
3
|
+
Dispatch `axstack-diligence` through Orca with a pinned brief and evidence paths.
|
|
4
|
+
It is read-only, never authors or edits, and returns `PASS` or `FINDINGS`
|
|
5
|
+
with locations, observed evidence, and limits. A stale or missing receipt is
|
|
6
|
+
not a pass. Keep its first pass independent of other reviewers and workers.
|
|
7
|
+
|
|
8
|
+
For a PR, compare every changed line with the accepted intent and exclusions:
|
|
9
|
+
is it intended and in scope? Check that no contract, rule, or obligation was
|
|
10
|
+
silently weakened or dropped by rewording. Compare the PR body, commit messages,
|
|
11
|
+
and author receipt with the diff: numbers, IDs, versions, test counts, sizes,
|
|
12
|
+
paths, and stale references. Bind the result to the exact head and base.
|
|
13
|
+
|
|
14
|
+
For research, reopen cited sources for answer-changing claims before the
|
|
15
|
+
driver folds verified claims. For a draft spec, compare it with Align decisions
|
|
16
|
+
before user approval: flag anything dropped, added, or softened. For tickets,
|
|
17
|
+
map every spec acceptance item to a capability's acceptance. Before candidate
|
|
18
|
+
publication, compare the author receipt with its evidence folder: red/green
|
|
19
|
+
logs exist, and counts, SHAs, and paths match. For release preparation, compare
|
|
20
|
+
the release PR body with the merged PRs.
|
|
21
|
+
|
|
22
|
+
`FINDINGS` identifies a mismatch for the driver to resolve at the owning phase;
|
|
23
|
+
it does not edit the artifact or create another review round by itself.
|
|
@@ -107,8 +107,8 @@ Tracking grants no merge, release, model-substitution, or scope authority.
|
|
|
107
107
|
|
|
108
108
|
The default 24-hour deadline covers standalone task-owned timers. Stop them at
|
|
109
109
|
deadline and preserve remaining work; the review automation has no task-owned
|
|
110
|
-
deadline.
|
|
111
|
-
|
|
110
|
+
deadline. Merge-ready requires applicable review receipt(s) and current diligence
|
|
111
|
+
`PASS` at the exact head; CI/tests alone are insufficient. Merge-ready
|
|
112
112
|
differs from merged; human merges.
|
|
113
113
|
|
|
114
114
|
## Review automation health
|
|
@@ -38,7 +38,7 @@ Read `roles.json` from the installed shared root `skills/axstack/`. The installe
|
|
|
38
38
|
shape is `{ "version": 1, "preset": "<name>", "roles": [...] }`. Bundled
|
|
39
39
|
profiles are setup inputs shaped as
|
|
40
40
|
`{ "version": 1, "roles": [...] }`. A new run records the selected preset and
|
|
41
|
-
all
|
|
41
|
+
all 32 role rows once. An active run keeps the exact snapshot until the user
|
|
42
42
|
explicitly changes it.
|
|
43
43
|
|
|
44
44
|
Select the requested role by stable ID. A missing or null model holds only that role;
|
|
@@ -0,0 +1,37 @@
|
|
|
1
|
+
# Role roster
|
|
2
|
+
|
|
3
|
+
- Chat drives (no role ID); `axstack-owner` owns one PR and
|
|
4
|
+
`axstack-author` its sole writer.
|
|
5
|
+
- `axstack-reviewer-primary` and `axstack-reviewer-secondary` are the ordered
|
|
6
|
+
peer pair. Peer review uses both; authored review uses this table:
|
|
7
|
+
|
|
8
|
+
| Preset | Author | Reviewer (model/effort) |
|
|
9
|
+
| --- | --- | --- |
|
|
10
|
+
| `mixed` | Codex / Sol (`codex/gpt-6-sol`) | `axstack-reviewer-secondary` (`claude/claude-opus-5-5` medium) |
|
|
11
|
+
| `mixed` | Claude / Opus (`claude/claude-opus-5-5`) | `axstack-reviewer-primary` (`codex/gpt-6-sol` high) |
|
|
12
|
+
| `codex-only` | Codex / Sol (`codex/gpt-6-sol`) | `axstack-reviewer-secondary` (`codex/gpt-6-luna` xhigh) |
|
|
13
|
+
| `claude-only` | Claude / Opus (`claude/claude-opus-5-5`) | `axstack-reviewer-secondary` (`claude/claude-sonnet-5-5` high) |
|
|
14
|
+
- `axstack-advisor-astra`/`axstack-advisor-opus` advise and author candidates;
|
|
15
|
+
`axstack-arena-candidate-grok`/
|
|
16
|
+
`axstack-arena-candidate-antigravity` add families.
|
|
17
|
+
`axstack-arena-judge-opus` judges round 1; `axstack-escalation-fable`/`axstack-arena-judge-astra` judge round 2.
|
|
18
|
+
High-stakes/trigger: fresh [contract](contracts.md) session.
|
|
19
|
+
`axstack-auditor` audits; `axstack-checker` reports discrepancies.
|
|
20
|
+
- `axstack-explainer`/`axstack-explainer-review`: explain/review.
|
|
21
|
+
- `axstack-diligence`: read-only [diligence checks](diligence.md) for every PR
|
|
22
|
+
review round and bounded research, spec, ticket, receipt, and release claims.
|
|
23
|
+
- `axstack-ui-verifier`: [UI checks](ui-verification.md).
|
|
24
|
+
- `axstack-auditor`/`axstack-research-requirements`/
|
|
25
|
+
`axstack-research-code`/`axstack-research-web`/
|
|
26
|
+
`axstack-explore-execution`/`axstack-monitor`:
|
|
27
|
+
`claude-sonnet-5-5` high in mixed/claude-only.
|
|
28
|
+
`axstack-monitor`: standalone watch never sends; chat-run watch: bounded
|
|
29
|
+
internal reports to its Run and original driver.
|
|
30
|
+
- Sol pairs `axstack-auditor-sol`/`axstack-research-code-sol`/
|
|
31
|
+
`axstack-explore-execution-sol`: `codex/gpt-6-sol` high in
|
|
32
|
+
mixed/codex-only; intentionally absent in claude-only. Dispatch each
|
|
33
|
+
independently from its Sonnet seat on the same bounded brief without
|
|
34
|
+
cross-reading. The driver reconciles findings per claim, never averages.
|
|
35
|
+
Record intentional absence and continue with Sonnet alone; a configured
|
|
36
|
+
but unavailable seat holds only its affected work.
|
|
37
|
+
- `axstack-debug-investigator-1..4` probe L1 briefs.
|
|
@@ -11,10 +11,10 @@ skills root, or an explicit user selection in the run record. Missing or contrad
|
|
|
11
11
|
a setup gap: hold. Never infer from live profiles or `list_profiles`, harness,
|
|
12
12
|
tools, credentials, quota, subscription, or default to `mixed`.
|
|
13
13
|
|
|
14
|
-
At start, snapshot all
|
|
14
|
+
At start, snapshot all 32 role IDs with provider/model/mode/effort; absent
|
|
15
15
|
or unconfigured roles are recorded explicitly; never default.
|
|
16
|
-
Such a role holds only
|
|
17
|
-
|
|
16
|
+
Such a role holds only its work. Later installed or changed roles need an
|
|
17
|
+
explicit user decision to enter the snapshot. Live profiles
|
|
18
18
|
are authoritative at snapshot time and for availability; bundled presets are setup
|
|
19
19
|
inputs, not runtime proof.
|
|
20
20
|
|
|
@@ -24,27 +24,7 @@ revalidation. Unavailable models, efforts, roles, or overrides hold only affecte
|
|
|
24
24
|
work; no automatic fallback, quota routing, subscription inference, or silent
|
|
25
25
|
provider/model/effort substitution.
|
|
26
26
|
|
|
27
|
-
|
|
28
|
-
`axstack-author` its sole writer.
|
|
29
|
-
- `axstack-reviewer-primary` and `axstack-reviewer-secondary` are the ordered
|
|
30
|
-
peer pair. Peer review uses both; authored review uses this table:
|
|
31
|
-
|
|
32
|
-
| Preset | Author | Reviewer (model/effort) |
|
|
33
|
-
| --- | --- | --- |
|
|
34
|
-
| `mixed` | Codex / Sol (`codex/gpt-6-sol`) | `axstack-reviewer-secondary` (`claude/claude-opus-5-5` medium) |
|
|
35
|
-
| `mixed` | Claude / Opus (`claude/claude-opus-5-5`) | `axstack-reviewer-primary` (`codex/gpt-6-sol` high) |
|
|
36
|
-
| `codex-only` | Codex / Sol (`codex/gpt-6-sol`) | `axstack-reviewer-secondary` (`codex/gpt-6-luna` xhigh) |
|
|
37
|
-
| `claude-only` | Claude / Opus (`claude/claude-opus-5-5`) | `axstack-reviewer-secondary` (`claude/claude-sonnet-5-5` high) |
|
|
38
|
-
- `axstack-advisor-astra`/`axstack-advisor-opus` advise and author candidates;
|
|
39
|
-
`axstack-arena-candidate-grok`/
|
|
40
|
-
`axstack-arena-candidate-antigravity` add families.
|
|
41
|
-
`axstack-arena-judge-opus` judges round 1; `axstack-escalation-fable`/`axstack-arena-judge-astra` judge round 2.
|
|
42
|
-
High-stakes/trigger: fresh [contract](contracts.md) session.
|
|
43
|
-
`axstack-auditor` audits; `axstack-checker` reports discrepancies.
|
|
44
|
-
- `axstack-explainer`/`axstack-explainer-review`: explain/review.
|
|
45
|
-
`axstack-monitor`: standalone watch never sends; chat-run watch: bounded
|
|
46
|
-
internal reports to its Run and original driver.
|
|
47
|
-
- `axstack-debug-investigator-1..4` probe L1 briefs.
|
|
27
|
+
Load the [Role roster](role-roster.md) for configured roles and authored-review pairings.
|
|
48
28
|
|
|
49
29
|
Provenance is matched on provider/model ID; effort never maps. Missing table-row
|
|
50
30
|
provenance is unsupported and `INCOMPLETE`; report it and ask the user. Never
|
|
@@ -55,18 +35,17 @@ step (3) for user routing: no substitution or same-provider review.
|
|
|
55
35
|
|
|
56
36
|
## Direct routes (no spec ceremony)
|
|
57
37
|
|
|
58
|
-
-
|
|
59
|
-
|
|
60
|
-
-
|
|
61
|
-
current/intended behavior
|
|
62
|
-
|
|
38
|
+
- Bounded research -> `axstack-research`: verify primary sources and code,
|
|
39
|
+
cite limits, and fan out distinct questions.
|
|
40
|
+
- Understand a system or gap -> `axstack-explain`:
|
|
41
|
+
current/intended behavior and bounded gaps from docs and renders.
|
|
42
|
+
"What could this break" follows
|
|
63
43
|
[Blast radius](blast-radius.md). Publication needs separate authority.
|
|
64
44
|
- A bug, failing test, regression, or wrong behavior, red loop wanted ->
|
|
65
45
|
`axstack-debug`: diagnose, escalate via adviser-directed investigators, hand
|
|
66
46
|
off a classified repair (explain: how; debug: what's wrong).
|
|
67
|
-
-
|
|
68
|
-
|
|
69
|
-
edits.
|
|
47
|
+
- Code quality/refactor discovery -> `axstack-improve`: rank bounded
|
|
48
|
+
candidates with evidence; report only, no source edits.
|
|
70
49
|
- Accepted worker/Task/Run completion or bounded backlog request -> driver invokes
|
|
71
50
|
`axstack-cleanup` inline; never dispatch it.
|
|
72
51
|
- Preparation completion, watch expiry, resume, or reconciliation -> the
|
|
@@ -0,0 +1,13 @@
|
|
|
1
|
+
# UI verification
|
|
2
|
+
|
|
3
|
+
Every Playwright, browser, or rendered-UI check, including a "confirm it in the
|
|
4
|
+
browser" step, goes through an Orca Dispatch to `axstack-ui-verifier` from the
|
|
5
|
+
run's role snapshot. Give it the exact build, URL, or artifact and the private
|
|
6
|
+
dispatch's evidence folder. The verifier is read-only: it never edits source.
|
|
7
|
+
The PR writer remains the sole writer.
|
|
8
|
+
|
|
9
|
+
Ask for screenshots and observed interactions, accessibility, desktop and
|
|
10
|
+
mobile layouts, and reduced-motion behavior where relevant. The verifier
|
|
11
|
+
returns a verdict with evidence paths and names checks it could not run.
|
|
12
|
+
Keep the verdict tied to the exact artifact or revision; changed bytes need a
|
|
13
|
+
fresh rendered pass.
|
|
@@ -28,8 +28,13 @@ immediately before an actual auditor profile or session dispatch. Ordinary
|
|
|
28
28
|
audit reading and record writing do not load it, and the auditor never
|
|
29
29
|
dispatches.
|
|
30
30
|
|
|
31
|
-
Core owns the `axstack-auditor` profile (
|
|
32
|
-
|
|
31
|
+
Core owns the `axstack-auditor` profile (claude/claude-sonnet-5-5 high in
|
|
32
|
+
mixed/claude-only; codex/gpt-6-luna xhigh in codex-only) and its invocation.
|
|
33
|
+
This skill governs what that auditor reads, measures, and proposes.
|
|
34
|
+
Dispatch `axstack-auditor` and `axstack-auditor-sol` independently on the same
|
|
35
|
+
bounded brief, without cross-reading. The driver reconciles findings per claim;
|
|
36
|
+
never average verdicts. Record an intentionally absent Sol seat and continue
|
|
37
|
+
with the base auditor alone; a configured but unavailable seat holds its work.
|
|
33
38
|
The user-chosen improvement mode is a tested, independently reviewed PR that a
|
|
34
39
|
human merges.
|
|
35
40
|
|
|
@@ -39,6 +39,8 @@ recorded reason.
|
|
|
39
39
|
loop cannot be built, stop, list what was tried, and ask the user for an
|
|
40
40
|
environment, a redacted artifact, or instrumentation permission. Done when
|
|
41
41
|
the command has run once and its red output is recorded.
|
|
42
|
+
Delegate any L0 or L1 headless-browser reproduction through
|
|
43
|
+
[UI verification](../axstack/references/ui-verification.md).
|
|
42
44
|
2. **Reproduce and minimise.** Confirm the loop reproduces the user's failure
|
|
43
45
|
and not a neighbour. Remove inputs, callers, config, data, and steps one at
|
|
44
46
|
a time within a stated budget until the repro is the smallest practical;
|
|
@@ -5,11 +5,14 @@ rendering matters.
|
|
|
5
5
|
|
|
6
6
|
1. Identify the final artifact bytes and theme. The explicit user theme wins;
|
|
7
7
|
otherwise use the dark default.
|
|
8
|
-
2.
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
8
|
+
2. Delegate the rendered pass through [UI verification](../../axstack/references/ui-verification.md).
|
|
9
|
+
Record actual desktop and mobile observations, or name the missing layout
|
|
10
|
+
check.
|
|
11
|
+
3. Have the verifier exercise relevant interactions, keyboard and screen-reader
|
|
12
|
+
accessibility, and reduced-motion behavior. Report each unavailable check
|
|
13
|
+
honestly.
|
|
14
|
+
4. The explainer reviewer checks text and source fidelity. Keep source
|
|
15
|
+
correctness, tests, rendered behavior, independent review, and publication
|
|
16
|
+
evidence separate.
|
|
14
17
|
5. Bind review to the exact artifact identity. Any byte change invalidates the
|
|
15
18
|
affected approval and requires fresh QA and review.
|
|
@@ -148,8 +148,9 @@ including evidence and retained complexity.
|
|
|
148
148
|
|
|
149
149
|
Run the acceptance checks and affected integration boundaries. Record commands,
|
|
150
150
|
observed outputs, and verified states. UI work includes rendered interaction
|
|
151
|
-
evidence
|
|
152
|
-
|
|
151
|
+
evidence through [UI verification](../axstack/references/ui-verification.md)
|
|
152
|
+
when relevant. Name every unavailable OS, harness, credential, or other
|
|
153
|
+
boundary instead of implying coverage.
|
|
153
154
|
|
|
154
155
|
After the last change, pin the exact candidate revision and return this compact
|
|
155
156
|
implementation receipt to the driver:
|
|
@@ -205,8 +206,13 @@ For each PR:
|
|
|
205
206
|
`REQUEST_CHANGES`, a failed required check, or post-readiness feedback returns
|
|
206
207
|
findings to the same author for a new revision, increments `repairs`, and
|
|
207
208
|
returns to step 1. `INCOMPLETE`, a provenance gap, unavailable model, serious
|
|
208
|
-
risk, or the third
|
|
209
|
-
parent sends its child back
|
|
209
|
+
risk, or the third review round with `REQUEST_CHANGES` and/or diligence
|
|
210
|
+
`FINDINGS` on one PR records `held`. A changed parent sends its child back
|
|
211
|
+
to step 1.
|
|
212
|
+
Merge-ready also requires a current diligence `PASS` at that head; diligence
|
|
213
|
+
`FINDINGS` return to the same author within the review round.
|
|
214
|
+
A round with reviewer `REQUEST_CHANGES` and/or diligence `FINDINGS` increments
|
|
215
|
+
`repairs` once and counts once toward the third-round hold.
|
|
210
216
|
|
|
211
217
|
One run-level completion wait covers every unsettled Dispatch; the bounded
|
|
212
218
|
forge check wait is the only other wait. End a turn only when every required PR
|
|
@@ -19,6 +19,11 @@ role dispatch, load the [Orca runtime
|
|
|
19
19
|
sequence](../axstack/references/orca-runtime.md). Use existing
|
|
20
20
|
`axstack-explore-codebase` or `axstack-research-code` roles only when their
|
|
21
21
|
specialization materially helps; create no new profile.
|
|
22
|
+
When dispatching `axstack-research-code`, dispatch `axstack-research-code-sol`
|
|
23
|
+
independently on the same bounded brief without cross-reading. The driver
|
|
24
|
+
reconciles findings per claim and never averages them. Record an intentionally
|
|
25
|
+
absent Sol pair and proceed with the base seat alone; a configured but
|
|
26
|
+
unavailable pair holds its work.
|
|
22
27
|
|
|
23
28
|
## 1. Bound discovery
|
|
24
29
|
|
|
@@ -29,8 +29,9 @@ is part of research.
|
|
|
29
29
|
|
|
30
30
|
2. **Fan out research:** A single factual lookup stays in the current chat.
|
|
31
31
|
Every other research run dispatches every configured research branch through
|
|
32
|
-
Orca: requirements
|
|
33
|
-
(Gemini/Antigravity, with Google Search
|
|
32
|
+
Orca: requirements, code, and web (Sonnet high in mixed/claude-only;
|
|
33
|
+
Codex in codex-only), web-google (Gemini/Antigravity, with Google Search
|
|
34
|
+
built in), and X (Grok).
|
|
34
35
|
Give each branch one owner, allow no cross-reading, and require a cited note
|
|
35
36
|
with a URL and access date per claim; re-open sources and never trust a search
|
|
36
37
|
summary. The driver reconciles agreements/disagreements per claim.
|
|
@@ -41,16 +42,27 @@ is part of research.
|
|
|
41
42
|
|
|
42
43
|
- `axstack-research-requirements`: requirements and intent.
|
|
43
44
|
- `axstack-research-code`: code behavior.
|
|
45
|
+
- `axstack-research-code-sol`: independent Sol code investigation.
|
|
44
46
|
- `axstack-research-web`: web and external sources.
|
|
45
47
|
- `axstack-research-web-google`: Google-Search-grounded web sources via Gemini/Antigravity.
|
|
46
48
|
- `axstack-research-x`: X (Twitter) posts and threads via Grok; cite post URLs and dates.
|
|
47
49
|
- `axstack-explore-codebase`: broad codebase mapping.
|
|
48
50
|
- `axstack-explore-execution`: execution and runtime traces.
|
|
51
|
+
- `axstack-explore-execution-sol`: independent Sol execution investigation.
|
|
52
|
+
|
|
53
|
+
When dispatching `axstack-research-code` or `axstack-explore-execution`,
|
|
54
|
+
dispatch its `-sol` pair independently on the same bounded brief without
|
|
55
|
+
cross-reading. The driver reconciles agreement and disagreement per claim,
|
|
56
|
+
never averaging findings. Record an intentionally absent pair and proceed
|
|
57
|
+
with the base seat alone; a configured but unavailable seat holds its work.
|
|
49
58
|
|
|
50
59
|
3. **Gather primary source evidence.** Inspect the actual documentation, code,
|
|
51
60
|
or tool output for every answer-changing claim. Apply the source standards
|
|
52
61
|
for citations, freshness, revisions, and access dates. Continue until each
|
|
53
62
|
material claim has direct evidence or a named evidence gap.
|
|
63
|
+
Before the driver folds verified claims, dispatch `axstack-diligence` under
|
|
64
|
+
[Diligence](../axstack/references/diligence.md) to reopen cited sources for
|
|
65
|
+
answer-changing claims and flag mismatches.
|
|
54
66
|
|
|
55
67
|
4. **Form the verdict.** Mark every material claim as **verified**,
|
|
56
68
|
**inference**, or **unverified** using the source standards. Derive
|
|
@@ -211,6 +211,8 @@ This section applies to peer and authored PR modes.
|
|
|
211
211
|
| `codex-only` | Codex / Sol (`codex/gpt-6-sol`) | `axstack-reviewer-secondary` (`codex/gpt-6-luna` xhigh) |
|
|
212
212
|
| `claude-only` | Claude / Opus (`claude/claude-opus-5-5`) | `axstack-reviewer-secondary` (`claude/claude-sonnet-5-5` high) |
|
|
213
213
|
|
|
214
|
+
The diligence receipt is separate and does not count as a reviewer receipt.
|
|
215
|
+
|
|
214
216
|
Provenance is matched on provider/model ID; record effort, but never use
|
|
215
217
|
effort to create a mapping. Any other author provenance for the
|
|
216
218
|
selected preset is unsupported and `INCOMPLETE`, including its secondary
|
|
@@ -234,6 +236,11 @@ This section applies to peer and authored PR modes.
|
|
|
234
236
|
|
|
235
237
|
Continue only when session receipts prove the required models, non-author
|
|
236
238
|
independence, actual author provenance where applicable, and exact brief.
|
|
239
|
+
In peer and authored PR review rounds, dispatch `axstack-diligence`
|
|
240
|
+
independently alongside the configured reviewer(s) on the same exact revision
|
|
241
|
+
and base. Follow [Diligence](../axstack/references/diligence.md). A diligence
|
|
242
|
+
`FINDINGS` receipt returns validated findings to the same author in the same
|
|
243
|
+
round; it is not an extra `REQUEST_CHANGES` round.
|
|
237
244
|
3. **Inspect all six angles.** In peer mode each reviewer covers every angle;
|
|
238
245
|
in authored mode the one reviewer covers all six angles:
|
|
239
246
|
1. Security and trust boundaries.
|
|
@@ -279,7 +286,8 @@ This section applies to peer and authored PR modes.
|
|
|
279
286
|
|
|
280
287
|
Verify the applicable spec, ticket, or intent acceptance, executable
|
|
281
288
|
evidence, exact candidate SHA, current base, and affected integration
|
|
282
|
-
boundary, plus rendered interaction evidence for relevant UI work
|
|
289
|
+
boundary, plus rendered interaction evidence for relevant UI work through
|
|
290
|
+
[UI verification](../axstack/references/ui-verification.md). A
|
|
283
291
|
passing test is insufficient when it checks the wrong behavior. Call out
|
|
284
292
|
seeded regressions, inadequate checks, and every unverified boundary. Every
|
|
285
293
|
mode-required receipt records concrete evidence and consequences, coverage,
|
|
@@ -320,12 +328,14 @@ These verdicts apply only to PR modes. Codebase findings use coverage status.
|
|
|
320
328
|
covering the whole brief, all six angles, and applicable acceptance.
|
|
321
329
|
- **Complete verdict:** validated blocking defects permit `REQUEST_CHANGES`;
|
|
322
330
|
complete evidence with no blocker permits `APPROVE`.
|
|
331
|
+
- **Diligence:** a current `PASS` is required with reviewer approval for
|
|
332
|
+
merge-ready; `FINDINGS` return to the author in that round.
|
|
323
333
|
- **Incomplete or stale:** use `INCOMPLETE`; never fabricate `APPROVE` or
|
|
324
334
|
`REQUEST_CHANGES`.
|
|
325
335
|
|
|
326
336
|
The owner verifies and synthesizes the mode-required evidence without voting.
|
|
327
|
-
A peer receipt count of one is incomplete; an authored
|
|
328
|
-
than one is not the selected mode. Passing tests or reviewer unanimity grants
|
|
337
|
+
A peer reviewer receipt count of one is incomplete; an authored reviewer
|
|
338
|
+
receipt count other than one is not the selected mode. Passing tests or reviewer unanimity grants
|
|
329
339
|
no merge authority.
|
|
330
340
|
|
|
331
341
|
## Template: candidate review brief
|
|
@@ -396,10 +406,11 @@ Mode-required exact-revision completeness gates external approval,
|
|
|
396
406
|
[merge-ready declarations](#authored-mode-own-pr), and authorized submission. It never gates returning
|
|
397
407
|
evidence, limitations, validated risk, or an internal `INCOMPLETE` report.
|
|
398
408
|
|
|
399
|
-
- Peer mode requires both current reviews
|
|
400
|
-
beyond the validated defects reported by
|
|
401
|
-
|
|
402
|
-
|
|
409
|
+
- Peer mode requires both current reviews, a separate current diligence receipt,
|
|
410
|
+
and no unresolved material finding beyond the validated defects reported by
|
|
411
|
+
`REQUEST_CHANGES`.
|
|
412
|
+
- Authored mode requires its one current eligible configured reviewer receipt and
|
|
413
|
+
a separate current diligence receipt; applicable scope identity must remain valid.
|
|
403
414
|
- A missing, mismatched, stale, or materially changed input blocks approval and
|
|
404
415
|
merge-ready declarations while readonly investigation continues.
|
|
405
416
|
|
|
@@ -54,6 +54,9 @@ and the lifecycle's [audit skill](../axstack-audit/SKILL.md) hook.
|
|
|
54
54
|
session to return plain AGREE. Present one
|
|
55
55
|
reviewable, identified revision for this checkpoint. Its user approval
|
|
56
56
|
creates the execution baseline.
|
|
57
|
+
Before user approval, dispatch `axstack-diligence` under
|
|
58
|
+
[Diligence](../axstack/references/diligence.md) to check the draft against
|
|
59
|
+
the Align decisions for anything dropped, added, or softened.
|
|
57
60
|
5. **Snapshot the baseline.** Record the approved revision identity and a
|
|
58
61
|
concise repository Markdown counterpart. In Linear mode, the native
|
|
59
62
|
document remains authoritative. In GitHub mode, the approved issue body is
|
|
@@ -54,6 +54,9 @@ an actual checker dispatch, not for ordinary mapping or state reconciliation.
|
|
|
54
54
|
rationale with the task; actual measurement and exception evidence follow in
|
|
55
55
|
the implement receipt. Mapping time requires no actual SHAs or line counts.
|
|
56
56
|
Every capability ends with the fields below and an explicit dependency list.
|
|
57
|
+
Before accepting the map, dispatch `axstack-diligence` under
|
|
58
|
+
[Diligence](../axstack/references/diligence.md) to confirm every spec
|
|
59
|
+
acceptance item maps to a capability's acceptance.
|
|
57
60
|
These routine mapping and split choices are autonomous driver decisions
|
|
58
61
|
within the approved spec; size alone never requires user approval.
|
|
59
62
|
|
|
@@ -135,6 +135,7 @@ The owner checks current required checks, all feedback, approvals, mergeability,
|
|
|
135
135
|
and exact-revision receipts before any merge-ready statement. API errors leave
|
|
136
136
|
readiness `UNKNOWN`; review approval alone is not merge-ready. Merge-ready is an
|
|
137
137
|
observed state distinct from merged, and the human merges by default.
|
|
138
|
+
A current diligence `PASS` at the exact head is required before any merge-ready statement.
|
|
138
139
|
Under authorized own-PR maintenance, keep repairing and rebasing onto the base
|
|
139
140
|
when it moves, then re-run checks, until the head is rebased on the current base,
|
|
140
141
|
every review comment and thread is addressed, at least one human team member's
|
package/src/roles.js
CHANGED
|
@@ -69,7 +69,7 @@ export function assessRoleReadiness(roles, preset) {
|
|
|
69
69
|
(preset === 'mixed' && role.id === 'axstack-arena-candidate-grok' && role.provider === 'grok') ||
|
|
70
70
|
(preset === 'mixed' && role.id === 'axstack-arena-candidate-antigravity' && role.provider === 'antigravity') ||
|
|
71
71
|
(preset === 'codex-only' && ['axstack-advisor-opus', 'axstack-escalation-fable', 'axstack-arena-judge-opus', 'axstack-research-web-google', 'axstack-research-x', 'axstack-arena-candidate-grok', 'axstack-arena-candidate-antigravity'].includes(role.id)) ||
|
|
72
|
-
(preset === 'claude-only' && ['axstack-advisor-astra', 'axstack-arena-judge-astra', 'axstack-research-web-google', 'axstack-research-x', 'axstack-arena-candidate-grok', 'axstack-arena-candidate-antigravity'].includes(role.id))
|
|
72
|
+
(preset === 'claude-only' && ['axstack-advisor-astra', 'axstack-arena-judge-astra', 'axstack-research-web-google', 'axstack-research-x', 'axstack-arena-candidate-grok', 'axstack-arena-candidate-antigravity', 'axstack-auditor-sol', 'axstack-research-code-sol', 'axstack-explore-execution-sol'].includes(role.id))
|
|
73
73
|
);
|
|
74
74
|
for (const role of roles) {
|
|
75
75
|
if (!bounds.has(role.provider)) {
|