axstack 0.20.30 → 0.20.31

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -104,11 +104,11 @@ Choose one explicit preset (32 roles each): [mixed](profiles/presets/mixed.json)
104
104
  (recommended), [codex-only](profiles/presets/codex-only.json), or
105
105
  [claude-only](profiles/presets/claude-only.json). Mixed supports cross-provider
106
106
  implementation review; single-provider presets have workflow limits and are
107
- not automatic fallbacks when a model is unavailable. See
107
+ not automatic cross-class or cross-provider fallbacks when a model is unavailable. See
108
108
  [workflow and routing details](docs/workflows.md).
109
109
 
110
110
  In `mixed` and `claude-only`, auditing, requirements/code/web research,
111
- execution exploration, and the optional monitor use Claude Sonnet 5.5 high.
111
+ execution exploration, and the optional monitor use the Claude Sonnet class at high effort.
112
112
  Auditing, code research, and execution exploration have independent Sol high
113
113
  pair seats in `mixed` and `codex-only`; `claude-only` records them as absent.
114
114
 
package/bin/axstack.js CHANGED
@@ -337,6 +337,7 @@ async function main() {
337
337
  if (summary.updated.length) console.log(`updated: ${summary.updated.join(', ')}`);
338
338
  if (summary.addedRoleIds?.length) console.log(`added role IDs: ${summary.addedRoleIds.join(', ')}`);
339
339
  if (summary.removedRoleIds?.length) console.log(`removed role IDs: ${summary.removedRoleIds.join(', ')}`);
340
+ if (summary.changedRoleModels?.length) console.log(`changed role models: ${summary.changedRoleModels.join(', ')}`);
340
341
  if (summary.removed.length) console.log(`removed: ${summary.removed.join(', ')}`);
341
342
  if (summary.unchanged.length) console.log(`unchanged: ${summary.unchanged.join(', ')}`);
342
343
  }
@@ -172,13 +172,14 @@ to rewrite them.
172
172
  ## Role behavior after installation
173
173
 
174
174
  The runtime reads `roles.json` from the installed shared root `skills/axstack/`.
175
- A new run records the selected preset plus all 32 role rows. An active run keeps
176
- that snapshot after a later preset install unless the user explicitly changes
177
- it and accepts the resulting evidence invalidation.
175
+ A new run records the selected preset plus all 32 role rows. Class rows resolve
176
+ at run start; each role snapshot records class, exact ID, source, and time.
177
+ An active run and resume reuse that snapshot after a later preset install unless
178
+ the user explicitly changes it and accepts the resulting evidence invalidation.
178
179
 
179
180
  The `mixed` and `claude-only` presets assign `axstack-auditor`,
180
181
  `axstack-research-requirements`, `axstack-research-code`, `axstack-research-web`,
181
- `axstack-explore-execution`, and `axstack-monitor` to Claude Sonnet 5.5 high.
182
+ `axstack-explore-execution`, and `axstack-monitor` to the Claude Sonnet class at high effort.
182
183
  The `codex-only` assignments for these roles are unchanged.
183
184
  The three `-sol` pair seats for auditor, research-code, and explore-execution
184
185
  use Sol high in `mixed` and `codex-only`; `claude-only` records intentional
@@ -198,8 +199,13 @@ receipts. For an arena-grade Align question, round 1 needs Opus; round 2, if
198
199
  invoked, needs escalation Fable and Astra; a required seat that is unavailable holds that
199
200
  round. The current chat drives on whatever
200
201
  model runs it; no preset carries a driver role. Every other missing, invalid, unsupported, or unavailable role value holds only
201
- the affected work. There is no model substitution, subscription inference, or
202
- quota routing.
202
+ the affected work. Codex class resolution reads the explicit catalog path via
203
+ `skills/axstack/scripts/resolve-models.js`; a missing or malformed catalog
204
+ holds. Claude launches an alias once per class, reads the exact ID from the
205
+ first assistant transcript turn, then reuses it. Only explicit model rejection
206
+ before that turn permits a recorded Codex retry within the same class, provider,
207
+ and effort. Claude rejection, timeout, quota, and auth failures hold; no
208
+ subscription inference or quota routing applies.
203
209
 
204
210
  `modeId` and similar permission fields remain conservative declared intent.
205
211
  They do not prove the effective Orca launcher mode, sandboxing, or permission
package/docs/workflows.md CHANGED
@@ -61,7 +61,7 @@ role.
61
61
 
62
62
  In `mixed` and `claude-only`, `axstack-auditor`, `axstack-research-requirements`,
63
63
  `axstack-research-code`, `axstack-research-web`, `axstack-explore-execution`,
64
- and `axstack-monitor` use Claude Sonnet 5.5 high. `codex-only` keeps its Codex
64
+ and `axstack-monitor` use the Claude Sonnet class at high effort. `codex-only` keeps its Codex
65
65
  assignments for those roles. Mixed web-google and X retain their source-specific
66
66
  Antigravity and Grok routes.
67
67
  The new `-sol` auditor, research-code, and explore-execution seats use Sol high
@@ -72,10 +72,15 @@ their findings per claim without averaging.
72
72
  The installed `<skills-dir>/axstack/roles.json` adds the selected preset name:
73
73
  `{ "version": 1, "preset": "<name>", "roles": [...] }`. The runtime reads it
74
74
  from the installed shared root `skills/axstack/` and records the whole table for
75
- a new run. Active runs retain their snapshot after later installation changes.
75
+ a new run. Per role it records class, exact ID, source, and time. Codex classes
76
+ resolve from a passed catalog path using `skills/axstack/scripts/resolve-models.js`;
77
+ missing or malformed catalogs hold. Claude's first class launch passes the alias,
78
+ then the first assistant transcript turn supplies the exact ID for later launches.
79
+ Unknown Claude IDs hold provenance-dependent work. Active runs and resume reuse
80
+ their snapshot after later installation changes.
76
81
 
77
82
  Peer roles keep the stable IDs `axstack-reviewer-primary` and
78
- `axstack-reviewer-secondary`; their provider/model mappings come only from the
83
+ `axstack-reviewer-secondary`; their provider/class mappings come only from the
79
84
  selected preset.
80
85
 
81
86
  The unavailable adviser in each single-provider preset stays explicitly
@@ -88,7 +93,10 @@ launchable; the run record snapshots the model the TUI reports. Missing or unava
88
93
  work. Model, effort, and permission values express requested intent until real
89
94
  Orca receipts establish the effective session. Stored `modeId` is not permission
90
95
  parity or a sandbox. No route is inferred from subscription, quota, harness,
91
- provider defaults, or installed tools, and no model is substituted silently.
96
+ provider defaults, or installed tools. Only explicit model rejection before the
97
+ first turn permits a recorded Codex retry with `--retry-of` to the next eligible
98
+ model in the same class, provider, and effort. Claude rejection, timeout, quota,
99
+ and auth failures hold.
92
100
 
93
101
  ## Orca runtime boundary
94
102
 
@@ -201,11 +209,15 @@ keeps the current owner and a resumable record.
201
209
 
202
210
  Serious security, downtime, data-loss, and major-design risks are raised in a
203
211
  prompt immediately and hold dependent dangerous work. This is not a runtime
204
- gate. An applicable `Notification policy` may use `axstack-relay` for serious
205
- risk immediately or a genuine blocker needing user intervention after bounded
206
- safe recovery. Questions, spec approvals, progress, CI pending, merge-ready,
207
- merged, and completion stay in Orca. The relay normally delivers one-way
208
- through native `hermes send`: it checks CLI lookup and the configured target,
212
+ gate. An applicable `Notification policy` may use `axstack-relay` only for a
213
+ user-decision hold (including spec or npm approval and a genuine blocker after
214
+ bounded safe recovery), a serious-risk hold immediately, or at most two merge-ready/merged milestones per run.
215
+ Routine questions stay in Orca. Progress, CI pending, and completion always stay
216
+ in Orca.
217
+ Only the bounded categories—user-decision holds (including spec approval),
218
+ serious-risk holds, and at most two merge-ready/merged milestones per run—may
219
+ be relayed under the recorded Notification policy. The relay normally delivers
220
+ one-way through native `hermes send`: it checks CLI lookup and the configured target,
209
221
  binds the recipient, deduplicates on the run record, and records the returned
210
222
  `message_id`. PR-manager notifications point the user to GitHub or a durable
211
223
  user-owned conversation; Telegram delivery, replies, and silence grant no action
@@ -216,9 +228,16 @@ read-only observer for standalone watches and never sends.
216
228
 
217
229
  ## Chat-run PR watch
218
230
 
231
+ For authorized engineering delivery, [Autopilot](../skills/axstack/references/autopilot.md)
232
+ continues from Align through the eligible phase sequence in the same chat.
233
+ The human approves substantial specs, every merge including release PRs, and
234
+ the npm stage. An open hold pauses the run. Implement arms maintain-mode watch
235
+ at its first published PR; release and install run only under recorded per-run
236
+ authority, and Close-out follows their verified receipts.
237
+
219
238
  Use `axstack-watch` chat-run mode to watch every PR raised by this chat's Run, including later verified publications and PRs the driver explicitly adopts. A harness-native monitoring or scheduled wake resumes the driver chat every 10 minutes by default; only when the harness has no such capability does the existing Orca `*/10` observer act as fallback. Record the chosen mechanism in the run record. Each wake runs the own-PR maintenance loop: address feedback, rebase on base movement, rerun required CI, and check the forge-counted human approval. Delegated work still runs through Orca; there is no daemon or polling model between wakes. Independent PRs can repair in parallel with one writer per PR; stack ancestor changes invalidate child evidence. An incomplete scan leaves readiness `UNKNOWN`.
220
239
 
221
- The watch lasts until all member PRs merge or close, you cancel it, or its wake expires. Stop and verify the chosen wake; an Orca fallback also needs automation disable/readback and workspace retirement. Worker settlement and run archive are separate driver steps. Run-created implementation candidates are published and read back before independent authored review. Adopted own-PR maintenance candidates receive independent exact-local-SHA review before driver publication and remote readback. The human merges. Source and installed instructions do not prove scheduled observation, driver wake, or live activation; those require native receipts.
240
+ The watch lasts until all member PRs merge or close and the run's release step is settled or not applicable, you cancel it, or its wake expires. Stop and verify the chosen wake; an Orca fallback also needs automation disable/readback and workspace retirement. Worker settlement and run archive are separate driver steps. Run-created implementation candidates are published and read back before independent authored review. Adopted own-PR maintenance candidates receive independent exact-local-SHA review before driver publication and remote readback. The human merges. Source and installed instructions do not prove scheduled observation, driver wake, or live activation; those require native receipts.
222
241
 
223
242
  ## Optional native peer-review automation
224
243
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "axstack",
3
- "version": "0.20.30",
3
+ "version": "0.20.31",
4
4
  "description": "Axstack installer and setup CLI: installs owned chat skills and role data, configures supported harness settings, and checks Orca capabilities.",
5
5
  "keywords": [
6
6
  "claude-code",
@@ -14,82 +14,82 @@
14
14
  "id": "axstack-advisor-opus",
15
15
  "name": "Axstack Opus adviser",
16
16
  "provider": "claude",
17
- "model": "claude-opus-5-5",
18
17
  "modeId": "bypassPermissions",
19
18
  "thinkingOptionId": "xhigh",
20
- "notes": "Independent Opus adviser for Align, Spec, and debug L1; authors the Claude arena candidate. Same bounded evidence and question as Astra; reuse only unchanged receipts."
19
+ "notes": "Independent Opus adviser for Align, Spec, and debug L1; authors the Claude arena candidate. Same bounded evidence and question as Astra; reuse only unchanged receipts. Claude alias resolves at first launch; a Claude rejection holds.",
20
+ "modelClass": "opus"
21
21
  },
22
22
  {
23
23
  "id": "axstack-owner",
24
24
  "name": "Axstack PR owner",
25
25
  "provider": "claude",
26
- "model": "claude-opus-5-5",
27
26
  "modeId": "bypassPermissions",
28
27
  "thinkingOptionId": "medium",
29
- "notes": "Persistent PR owner: one owner per PR, accountable for candidate, fixes, verification evidence, and monitoring. May delegate coding but never edits a worker-owned candidate concurrently. Launches eligible independent reviewers."
28
+ "notes": "Persistent PR owner: one owner per PR, accountable for candidate, fixes, verification evidence, and monitoring. May delegate coding but never edits a worker-owned candidate concurrently. Launches eligible independent reviewers. Claude alias resolves at first launch; a Claude rejection holds.",
29
+ "modelClass": "opus"
30
30
  },
31
31
  {
32
32
  "id": "axstack-author",
33
33
  "name": "Axstack author",
34
34
  "provider": "claude",
35
- "model": "claude-opus-5-5",
36
35
  "modeId": "bypassPermissions",
37
36
  "thinkingOptionId": "medium",
38
- "notes": "Ordinary implementation and repairs. Uses strict red-green-refactor and remains the exclusive writer for a candidate. Validate configured availability at launch; hold affected work without fallback."
37
+ "notes": "Ordinary implementation and repairs. Uses strict red-green-refactor and remains the exclusive writer for a candidate. Claude alias resolves at first launch; a Claude rejection holds.",
38
+ "modelClass": "opus"
39
39
  },
40
40
  {
41
41
  "id": "axstack-reviewer-primary",
42
42
  "name": "Axstack reviewer (primary)",
43
43
  "provider": "claude",
44
- "model": "claude-opus-5-5",
45
44
  "modeId": "bypassPermissions",
46
45
  "thinkingOptionId": "medium",
47
- "notes": "Primary reviewer in the ordered claude-only peer pair: Opus medium followed by Sonnet high. Never author or owner; review exact SHA and base across all six angles and acceptance. Peer first pass stays isolated."
46
+ "notes": "Primary reviewer in the ordered claude-only peer pair: Opus medium followed by Sonnet high. Never author or owner; review exact SHA and base across all six angles and acceptance. Peer first pass stays isolated. Claude alias resolves at first launch; a Claude rejection holds.",
47
+ "modelClass": "opus"
48
48
  },
49
49
  {
50
50
  "id": "axstack-reviewer-secondary",
51
51
  "name": "Axstack reviewer (secondary)",
52
52
  "provider": "claude",
53
- "model": "claude-sonnet-5-5",
54
53
  "modeId": "bypassPermissions",
55
54
  "thinkingOptionId": "high",
56
- "notes": "Secondary reviewer in the ordered claude-only peer pair: Opus medium followed by Sonnet high. Eligible authored reviewer for an Opus-authored candidate. Never author or owner; review exact SHA and base across all six angles and acceptance. Peer first pass stays isolated."
55
+ "notes": "Secondary reviewer in the ordered claude-only peer pair: Opus medium followed by Sonnet high. Eligible authored reviewer for an Opus-authored candidate. Never author or owner; review exact SHA and base across all six angles and acceptance. Peer first pass stays isolated. Claude alias resolves at first launch; a Claude rejection holds.",
56
+ "modelClass": "sonnet"
57
57
  },
58
58
  {
59
59
  "id": "axstack-diligence",
60
60
  "name": "Axstack diligence checker",
61
61
  "provider": "claude",
62
- "model": "claude-sonnet-5-5",
63
62
  "modeId": "bypassPermissions",
64
63
  "thinkingOptionId": "high",
65
- "notes": "Read-only diligence for exact-revision PRs and bounded research, spec, ticket, receipt, and release claims. Returns PASS or FINDINGS with evidence; never authors or edits."
64
+ "notes": "Read-only diligence for exact-revision PRs and bounded research, spec, ticket, receipt, and release claims. Returns PASS or FINDINGS with evidence; never authors or edits. Claude alias resolves at first launch; a Claude rejection holds.",
65
+ "modelClass": "sonnet"
66
66
  },
67
67
  {
68
68
  "id": "axstack-checker",
69
69
  "name": "Axstack tracker checker",
70
70
  "provider": "claude",
71
- "model": "claude-sonnet-5-5",
72
71
  "modeId": "bypassPermissions",
73
72
  "thinkingOptionId": "high",
74
- "notes": "Report-only discrepancy checker for the selected external tracker (Linear or GitHub Issues). Never mutates the tracker; the driver independently verifies evidence before applying updates. A null model means explicit user selection is required before dispatch and must never launch a provider default."
73
+ "notes": "Report-only discrepancy checker for the selected external tracker (Linear or GitHub Issues). Never mutates the tracker; the driver independently verifies evidence before applying updates. A null model means explicit user selection is required before dispatch and must never launch a provider default. Claude alias resolves at first launch; a Claude rejection holds.",
74
+ "modelClass": "sonnet"
75
75
  },
76
76
  {
77
77
  "id": "axstack-research-requirements",
78
78
  "name": "Axstack research requirements",
79
79
  "provider": "claude",
80
- "model": "claude-sonnet-5-5",
81
80
  "modeId": "bypassPermissions",
82
81
  "thinkingOptionId": "high",
83
- "notes": "Sonnet high research requirements analyst: scopes bounded questions and acceptance for a research task. Validate configured availability at launch; hold affected work without fallback."
82
+ "notes": "Sonnet high research requirements analyst: scopes bounded questions and acceptance for a research task. Claude alias resolves at first launch; a Claude rejection holds.",
83
+ "modelClass": "sonnet"
84
84
  },
85
85
  {
86
86
  "id": "axstack-research-code",
87
87
  "name": "Axstack research code",
88
88
  "provider": "claude",
89
- "model": "claude-sonnet-5-5",
90
89
  "modeId": "bypassPermissions",
91
90
  "thinkingOptionId": "high",
92
- "notes": "Sonnet high research code investigator: verifies behavior against inspected code and executable evidence. Validate configured availability at launch; hold affected work without fallback."
91
+ "notes": "Sonnet high research code investigator: verifies behavior against inspected code and executable evidence. Claude alias resolves at first launch; a Claude rejection holds.",
92
+ "modelClass": "sonnet"
93
93
  },
94
94
  {
95
95
  "id": "axstack-research-code-sol",
@@ -104,10 +104,10 @@
104
104
  "id": "axstack-research-web",
105
105
  "name": "Axstack research web",
106
106
  "provider": "claude",
107
- "model": "claude-sonnet-5-5",
108
107
  "modeId": "bypassPermissions",
109
108
  "thinkingOptionId": "high",
110
- "notes": "Sonnet high research web reader: gathers primary-source facts efficiently. Validate configured availability at launch; hold affected work without fallback."
109
+ "notes": "Sonnet high research web reader: gathers primary-source facts efficiently. Claude alias resolves at first launch; a Claude rejection holds.",
110
+ "modelClass": "sonnet"
111
111
  },
112
112
  {
113
113
  "id": "axstack-research-web-google",
@@ -131,46 +131,46 @@
131
131
  "id": "axstack-explainer",
132
132
  "name": "Axstack explainer",
133
133
  "provider": "claude",
134
- "model": "claude-sonnet-5-5",
135
134
  "modeId": "bypassPermissions",
136
135
  "thinkingOptionId": "high",
137
- "notes": "Complex visual explanation author: traces systems, changes, and implementation gaps in requested artifacts and verifies rendered behavior where applicable. Validate configured availability at launch; hold affected work without fallback."
136
+ "notes": "Complex visual explanation author: traces systems, changes, and implementation gaps in requested artifacts and verifies rendered behavior where applicable. Claude alias resolves at first launch; a Claude rejection holds.",
137
+ "modelClass": "sonnet"
138
138
  },
139
139
  {
140
140
  "id": "axstack-explainer-review",
141
141
  "name": "Axstack explainer reviewer",
142
142
  "provider": "claude",
143
- "model": "claude-sonnet-5-5",
144
143
  "modeId": "bypassPermissions",
145
144
  "thinkingOptionId": "high",
146
- "notes": "Independent visual explanation reviewer: checks the exact artifact for text and source fidelity. The rendered pass belongs to axstack-ui-verifier. Any artifact change invalidates its review. Validate configured availability at launch; hold affected work without fallback."
145
+ "notes": "Independent visual explanation reviewer: checks the exact artifact for text and source fidelity. The rendered pass belongs to axstack-ui-verifier. Any artifact change invalidates its review. Claude alias resolves at first launch; a Claude rejection holds.",
146
+ "modelClass": "sonnet"
147
147
  },
148
148
  {
149
149
  "id": "axstack-ui-verifier",
150
150
  "name": "Axstack UI verifier",
151
151
  "provider": "claude",
152
- "model": "claude-sonnet-5-5",
153
152
  "modeId": "bypassPermissions",
154
153
  "thinkingOptionId": "high",
155
- "notes": "Read-only UI verifier: runs Playwright or the browser against the given build, URL, or artifact; captures screenshots, interactions, accessibility, desktop/mobile, and reduced-motion evidence in the dispatch evidence folder; returns a verdict with evidence paths. Never edits source. Validate configured availability at launch; hold affected work without fallback."
154
+ "notes": "Read-only UI verifier: runs Playwright or the browser against the given build, URL, or artifact; captures screenshots, interactions, accessibility, desktop/mobile, and reduced-motion evidence in the dispatch evidence folder; returns a verdict with evidence paths. Never edits source. Claude alias resolves at first launch; a Claude rejection holds.",
155
+ "modelClass": "sonnet"
156
156
  },
157
157
  {
158
158
  "id": "axstack-explore-codebase",
159
159
  "name": "Axstack codebase explorer",
160
160
  "provider": "claude",
161
- "model": "claude-sonnet-5-5",
162
161
  "modeId": "bypassPermissions",
163
162
  "thinkingOptionId": "high",
164
- "notes": "Codebase mapper: explores repository structure and interfaces for research and handoff context. Validate configured availability at launch; hold affected work without fallback."
163
+ "notes": "Codebase mapper: explores repository structure and interfaces for research and handoff context. Claude alias resolves at first launch; a Claude rejection holds.",
164
+ "modelClass": "sonnet"
165
165
  },
166
166
  {
167
167
  "id": "axstack-explore-execution",
168
168
  "name": "Axstack execution explorer",
169
169
  "provider": "claude",
170
- "model": "claude-sonnet-5-5",
171
170
  "modeId": "bypassPermissions",
172
171
  "thinkingOptionId": "high",
173
- "notes": "Sonnet high execution explorer: runs bounded checks of runtime behavior where authorized. Validate configured availability at launch; hold affected work without fallback."
172
+ "notes": "Sonnet high execution explorer: runs bounded checks of runtime behavior where authorized. Claude alias resolves at first launch; a Claude rejection holds.",
173
+ "modelClass": "sonnet"
174
174
  },
175
175
  {
176
176
  "id": "axstack-explore-execution-sol",
@@ -185,19 +185,19 @@
185
185
  "id": "axstack-monitor",
186
186
  "name": "Axstack monitor",
187
187
  "provider": "claude",
188
- "model": "claude-sonnet-5-5",
189
188
  "modeId": "bypassPermissions",
190
189
  "thinkingOptionId": "high",
191
- "notes": "Optional Sonnet high independent read-only observer for a standalone PR watch. Reads GitHub, feedback, and checks, persists event IDs, and wakes the owner only for a new actionable event. Never sends, authors, reviews, replies, or acts as either reusable PR manager. Healthy snapshots stay quiet. Chat-run mode: one same-host native read-only observer per Run; fresh finite passes report precise deltas internally to the original Run/driver and may disable/read back only their own automation at verified stop. No repair, dispatch, public notification, or replacement coordinator. Effective scheduled model/effort and wake require live proof."
190
+ "notes": "Optional Sonnet high independent read-only observer for a standalone PR watch. Reads GitHub, feedback, and checks, persists event IDs, and wakes the owner only for a new actionable event. Never sends, authors, reviews, replies, or acts as either reusable PR manager. Healthy snapshots stay quiet. Chat-run mode: one same-host native read-only observer per Run; fresh finite passes report precise deltas internally to the original Run/driver and may disable/read back only their own automation at verified stop. No repair, dispatch, public notification, or replacement coordinator. Effective scheduled model/effort and wake require live proof. Claude alias resolves at first launch; a Claude rejection holds.",
191
+ "modelClass": "sonnet"
192
192
  },
193
193
  {
194
194
  "id": "axstack-auditor",
195
195
  "name": "Axstack auditor",
196
196
  "provider": "claude",
197
- "model": "claude-sonnet-5-5",
198
197
  "modeId": "bypassPermissions",
199
198
  "thinkingOptionId": "high",
200
- "notes": "Sonnet high read-only end-of-run and checkpoint auditor. Collects scope and outcome evidence with counts and denominators and reports PASS, FAIL, or UNKNOWN without inventing numbers. Never edits, merges, activates, or audits itself."
199
+ "notes": "Sonnet high read-only end-of-run and checkpoint auditor. Collects scope and outcome evidence with counts and denominators and reports PASS, FAIL, or UNKNOWN without inventing numbers. Never edits, merges, activates, or audits itself. Claude alias resolves at first launch; a Claude rejection holds.",
200
+ "modelClass": "sonnet"
201
201
  },
202
202
  {
203
203
  "id": "axstack-auditor-sol",
@@ -212,37 +212,37 @@
212
212
  "id": "axstack-debug-investigator-1",
213
213
  "name": "Axstack debug investigator 1",
214
214
  "provider": "claude",
215
- "model": "claude-opus-5-5",
216
215
  "modeId": "bypassPermissions",
217
216
  "thinkingOptionId": "medium",
218
- "notes": "Debug investigator seat 1. Dispatched only by axstack-debug at L1 with the shared evidence packet and one distinct brief; never reads another investigator's output. Works in its own disposable worktree at the pinned revision plus the recorded dirty patch; may instrument there for probes; never commits, pushes, publishes, or creates children. Returns one receipt per brief. Independence comes from brief isolation, not model diversity. This preset repeats claude-opus-5-5 at medium effort because it has fewer model families."
217
+ "notes": "Debug investigator seat 1. Dispatched only by axstack-debug at L1 with the shared evidence packet and one distinct brief; never reads another investigator's output. Works in its own disposable worktree at the pinned revision plus the recorded dirty patch; may instrument there for probes; never commits, pushes, publishes, or creates children. Returns one receipt per brief. Independence comes from brief isolation, not model diversity. This preset repeats the Opus class at medium effort because it has fewer model families. Claude alias resolves at first launch; a Claude rejection holds.",
218
+ "modelClass": "opus"
219
219
  },
220
220
  {
221
221
  "id": "axstack-debug-investigator-2",
222
222
  "name": "Axstack debug investigator 2",
223
223
  "provider": "claude",
224
- "model": "claude-sonnet-5-5",
225
224
  "modeId": "bypassPermissions",
226
225
  "thinkingOptionId": "high",
227
- "notes": "Debug investigator seat 2. Dispatched only by axstack-debug at L1 with the shared evidence packet and one distinct brief; never reads another investigator's output. Works in its own disposable worktree at the pinned revision plus the recorded dirty patch; may instrument there for probes; never commits, pushes, publishes, or creates children. Returns one receipt per brief. Independence comes from brief isolation, not model diversity. This preset repeats claude-sonnet-5-5 at high effort because it has fewer model families; independence comes from brief isolation."
226
+ "notes": "Debug investigator seat 2. Dispatched only by axstack-debug at L1 with the shared evidence packet and one distinct brief; never reads another investigator's output. Works in its own disposable worktree at the pinned revision plus the recorded dirty patch; may instrument there for probes; never commits, pushes, publishes, or creates children. Returns one receipt per brief. Independence comes from brief isolation, not model diversity. This preset repeats the Sonnet class at high effort because it has fewer model families; independence comes from brief isolation. Claude alias resolves at first launch; a Claude rejection holds.",
227
+ "modelClass": "sonnet"
228
228
  },
229
229
  {
230
230
  "id": "axstack-debug-investigator-3",
231
231
  "name": "Axstack debug investigator 3",
232
232
  "provider": "claude",
233
- "model": "claude-opus-5-5",
234
233
  "modeId": "bypassPermissions",
235
234
  "thinkingOptionId": "medium",
236
- "notes": "Debug investigator seat 3. Dispatched only by axstack-debug at L1 with the shared evidence packet and one distinct brief; never reads another investigator's output. Works in its own disposable worktree at the pinned revision plus the recorded dirty patch; may instrument there for probes; never commits, pushes, publishes, or creates children. Returns one receipt per brief. Independence comes from brief isolation, not model diversity. This preset repeats claude-opus-5-5 at medium effort because it has fewer model families."
235
+ "notes": "Debug investigator seat 3. Dispatched only by axstack-debug at L1 with the shared evidence packet and one distinct brief; never reads another investigator's output. Works in its own disposable worktree at the pinned revision plus the recorded dirty patch; may instrument there for probes; never commits, pushes, publishes, or creates children. Returns one receipt per brief. Independence comes from brief isolation, not model diversity. This preset repeats the Opus class at medium effort because it has fewer model families. Claude alias resolves at first launch; a Claude rejection holds.",
236
+ "modelClass": "opus"
237
237
  },
238
238
  {
239
239
  "id": "axstack-debug-investigator-4",
240
240
  "name": "Axstack debug investigator 4",
241
241
  "provider": "claude",
242
- "model": "claude-sonnet-5-5",
243
242
  "modeId": "bypassPermissions",
244
243
  "thinkingOptionId": "high",
245
- "notes": "Debug investigator seat 4. Dispatched only by axstack-debug at L1 with the shared evidence packet and one distinct brief; never reads another investigator's output. Works in its own disposable worktree at the pinned revision plus the recorded dirty patch; may instrument there for probes; never commits, pushes, publishes, or creates children. Returns one receipt per brief. Independence comes from brief isolation, not model diversity. This preset repeats claude-sonnet-5-5 at high effort because it has fewer model families; independence comes from brief isolation."
244
+ "notes": "Debug investigator seat 4. Dispatched only by axstack-debug at L1 with the shared evidence packet and one distinct brief; never reads another investigator's output. Works in its own disposable worktree at the pinned revision plus the recorded dirty patch; may instrument there for probes; never commits, pushes, publishes, or creates children. Returns one receipt per brief. Independence comes from brief isolation, not model diversity. This preset repeats the Sonnet class at high effort because it has fewer model families; independence comes from brief isolation. Claude alias resolves at first launch; a Claude rejection holds.",
245
+ "modelClass": "sonnet"
246
246
  },
247
247
  {
248
248
  "id": "axstack-arena-judge-astra",
@@ -257,19 +257,19 @@
257
257
  "id": "axstack-escalation-fable",
258
258
  "name": "Axstack escalation Fable",
259
259
  "provider": "claude",
260
- "model": "claude-fable-5-1",
261
260
  "modeId": "bypassPermissions",
262
261
  "thinkingOptionId": "xhigh",
263
- "notes": "Fable escalation seat for arena round 2, high-stakes plain AGREE, or the bounded escalation trigger. Fresh session per use; never reuses adviser or candidate context. Read-only round 2 judge: scores every candidate by label; never authors a candidate."
262
+ "notes": "Fable escalation seat for arena round 2, high-stakes plain AGREE, or the bounded escalation trigger. Fresh session per use; never reuses adviser or candidate context. Read-only round 2 judge: scores every candidate by label; never authors a candidate. Claude alias resolves at first launch; a Claude rejection holds.",
263
+ "modelClass": "fable"
264
264
  },
265
265
  {
266
266
  "id": "axstack-arena-judge-opus",
267
267
  "name": "Axstack arena judge Opus",
268
268
  "provider": "claude",
269
- "model": "claude-opus-5-5",
270
269
  "modeId": "bypassPermissions",
271
270
  "thinkingOptionId": "xhigh",
272
- "notes": "Read-only round 1 arena judge. Receives the rubric and every candidate by label, scores each criterion, and recommends a base with rationale. Never authors a candidate, never cross-reads another judge, never mutates; the driver compares its verdict without averaging."
271
+ "notes": "Read-only round 1 arena judge. Receives the rubric and every candidate by label, scores each criterion, and recommends a base with rationale. Never authors a candidate, never cross-reads another judge, never mutates; the driver compares its verdict without averaging. Claude alias resolves at first launch; a Claude rejection holds.",
272
+ "modelClass": "opus"
273
273
  },
274
274
  {
275
275
  "id": "axstack-arena-candidate-grok",