axstack 0.13.1 → 0.14.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -77,7 +77,7 @@ targets are reported without normal-path adoption. Install exits nonzero when
77
77
  an instruction conflict is preserved, while clean and idempotent installs exit
78
78
  successfully.
79
79
 
80
- The public bundle preserves three canonical 21-role inputs:
80
+ The public bundle preserves three canonical 23-role inputs:
81
81
  [mixed](profiles/presets/mixed.json),
82
82
  [codex-only](profiles/presets/codex-only.json), and
83
83
  [claude-only](profiles/presets/claude-only.json). Each is exactly
@@ -51,7 +51,7 @@ profiles/presets/codex-only.json
51
51
  profiles/presets/claude-only.json
52
52
  ```
53
53
 
54
- Each has exactly `{ "version": 1, "roles": [...] }` with the same 21 stable
54
+ Each has exactly `{ "version": 1, "roles": [...] }` with the same 23 stable
55
55
  role IDs. Installation writes `<skills-dir>/axstack/roles.json` as
56
56
  `{ "version": 1, "preset": "<selected preset>", "roles": [...] }` and records
57
57
  its ownership hash like every other installed skill asset. There is no second
@@ -118,9 +118,9 @@ The complete bundle is validated before writes:
118
118
  - each preset is a real JSON file with version 1, a non-empty `roles` array,
119
119
  the filename's selected identity supplied by the caller, and the same role-ID
120
120
  set as its peers;
121
- - every role has valid preserved fields, while the mixed checker and the
122
- unavailable adviser in each single-provider preset explicitly permit
123
- `model: null`;
121
+ - every role has valid preserved fields, while the mixed checker and, in each
122
+ single-provider preset, the unavailable adviser and its matching arena judge
123
+ seat explicitly permit `model: null`;
124
124
  - obsolete runtime configuration flags fail before mutation with migration
125
125
  guidance.
126
126
 
@@ -148,15 +148,16 @@ to rewrite them.
148
148
  ## Role behavior after installation
149
149
 
150
150
  The runtime reads `roles.json` relative to the actually loaded `axstack` skill.
151
- A new run records the selected preset plus all 21 role rows. An active run keeps
151
+ A new run records the selected preset plus all 23 role rows. An active run keeps
152
152
  that snapshot after a later preset install unless the user explicitly changes
153
153
  it and accepts the resulting evidence invalidation.
154
154
 
155
155
  The mixed checker has `model: null`; checker dispatch is held and never inherits
156
156
  a provider default. The single-provider presets configure the checker. Their
157
- unavailable adviser remains an explicit same-provider `model: null` role, which
158
- does not make installation unready; Align and Spec still hold until both Astra
159
- and Fable can return independent receipts. The current chat drives on whatever
157
+ unavailable adviser and its matching arena judge seat remain explicit
158
+ same-provider `model: null` roles, which do not make installation unready;
159
+ Align and Spec still hold until both Astra and Fable can return independent
160
+ receipts, and an arena-grade Align question holds until both judge seats can. The current chat drives on whatever
160
161
  model runs it; no preset carries a driver role. Every other missing, invalid, unsupported, or unavailable role value holds only
161
162
  the affected work. There is no model substitution, subscription inference, or
162
163
  quota routing.
package/docs/workflows.md CHANGED
@@ -40,7 +40,7 @@ only affected work.
40
40
 
41
41
  Installation requires one explicit canonical preset. The three bundle files
42
42
  under `profiles/presets/` each contain exactly
43
- `{ "version": 1, "roles": [...] }` and the same 21 stable IDs.
43
+ `{ "version": 1, "roles": [...] }` and the same 23 stable IDs.
44
44
 
45
45
  The current chat drives on whatever model runs it; no preset carries a driver
46
46
  role.
@@ -108,7 +108,13 @@ session and evidence remain valid.
108
108
 
109
109
  - `axstack-align` maps facts and dependencies, asks prioritized questions, and
110
110
  consults Astra and Fable independently with the same bounded evidence and
111
- question. It synthesizes disagreements and reuses unchanged receipts.
111
+ question. It synthesizes disagreements and reuses unchanged receipts. For a
112
+ hard-to-reverse design choice it runs one arena round instead: Astra and
113
+ Fable each author a candidate, `axstack-arena-judge-astra` and
114
+ `axstack-arena-judge-fable` score both against the driver's rubric, and the
115
+ driver picks a base, grafts the losers' strong ideas, and presents the
116
+ synthesis as the recommendation; the note lands as `Decisions` rows in the
117
+ run record.
112
118
  - `axstack-spec` writes observable acceptance, exclusions, decisions, and one
113
119
  user-approved revision baseline.
114
120
  - `axstack-tickets` maps user-visible capabilities to dependency-aware internal
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "axstack",
3
- "version": "0.13.1",
3
+ "version": "0.14.1",
4
4
  "description": "Axstack installer and setup CLI: installs owned chat skills and role data, configures supported harness settings, and checks Orca capabilities.",
5
5
  "keywords": [
6
6
  "claude-code",
@@ -189,6 +189,24 @@
189
189
  "modeId": "bypassPermissions",
190
190
  "thinkingOptionId": "xhigh",
191
191
  "notes": "Debug investigator seat 4. Dispatched only by axstack-debug at L1 with the shared evidence packet and one distinct brief; never reads another investigator's output. Works in its own disposable worktree at the pinned revision plus the recorded dirty patch; may instrument there for probes; never commits, pushes, publishes, or creates children. Returns one receipt per brief. Independence comes from brief isolation, not model diversity. This preset repeats claude-sonnet-5 at xhigh effort because it has fewer model families; independence comes from brief isolation."
192
+ },
193
+ {
194
+ "id": "axstack-arena-judge-astra",
195
+ "name": "Axstack arena judge Astra",
196
+ "provider": "claude",
197
+ "model": null,
198
+ "modeId": "bypassPermissions",
199
+ "thinkingOptionId": "xhigh",
200
+ "notes": "Required Arena judge (Astra seat). Read-only cross-judge for an arena-grade Align decision: receives the rubric and both candidates by label, scores each criterion, and recommends a base with rationale. Never authors a candidate, never cross-reads the other judge, never mutates. Disagreement between judges is surfaced by the driver, never averaged. Intentionally absent in the claude-only preset; the arena-grade decision holds without substitution."
201
+ },
202
+ {
203
+ "id": "axstack-arena-judge-fable",
204
+ "name": "Axstack arena judge Fable",
205
+ "provider": "claude",
206
+ "model": "claude-fable-5-1",
207
+ "modeId": "bypassPermissions",
208
+ "thinkingOptionId": "xhigh",
209
+ "notes": "Arena judge (Fable seat). Read-only cross-judge for an arena-grade Align decision: receives the rubric and both candidates by label, scores each criterion, and recommends a base with rationale. Never authors a candidate, never cross-reads the other judge, never mutates. Disagreement between judges is surfaced by the driver, never averaged."
192
210
  }
193
211
  ]
194
212
  }
@@ -189,6 +189,24 @@
189
189
  "modeId": "full-access",
190
190
  "thinkingOptionId": "xhigh",
191
191
  "notes": "Debug investigator seat 4. Dispatched only by axstack-debug at L1 with the shared evidence packet and one distinct brief; never reads another investigator's output. Works in its own disposable worktree at the pinned revision plus the recorded dirty patch; may instrument there for probes; never commits, pushes, publishes, or creates children. Returns one receipt per brief. Independence comes from brief isolation, not model diversity. This preset repeats gpt-5.6-terra at xhigh effort because it has fewer model families."
192
+ },
193
+ {
194
+ "id": "axstack-arena-judge-astra",
195
+ "name": "Axstack arena judge Astra",
196
+ "provider": "codex",
197
+ "model": "gpt-6-astra",
198
+ "modeId": "full-access",
199
+ "thinkingOptionId": "xhigh",
200
+ "notes": "Arena judge (Astra seat). Read-only cross-judge for an arena-grade Align decision: receives the rubric and both candidates by label, scores each criterion, and recommends a base with rationale. Never authors a candidate, never cross-reads the other judge, never mutates. Disagreement between judges is surfaced by the driver, never averaged."
201
+ },
202
+ {
203
+ "id": "axstack-arena-judge-fable",
204
+ "name": "Axstack arena judge Fable",
205
+ "provider": "codex",
206
+ "model": null,
207
+ "modeId": "full-access",
208
+ "thinkingOptionId": "xhigh",
209
+ "notes": "Required Arena judge (Fable seat). Read-only cross-judge for an arena-grade Align decision: receives the rubric and both candidates by label, scores each criterion, and recommends a base with rationale. Never authors a candidate, never cross-reads the other judge, never mutates. Disagreement between judges is surfaced by the driver, never averaged. Intentionally absent in the codex-only preset; the arena-grade decision holds without substitution."
192
210
  }
193
211
  ]
194
212
  }
@@ -189,6 +189,24 @@
189
189
  "modeId": "full-access",
190
190
  "thinkingOptionId": "low",
191
191
  "notes": "Debug investigator seat 4. Dispatched only by axstack-debug at L1 with the shared evidence packet and one distinct brief; never reads another investigator's output. Works in its own disposable worktree at the pinned revision plus the recorded dirty patch; may instrument there for probes; never commits, pushes, publishes, or creates children. Returns one receipt per brief. Independence comes from brief isolation, not model diversity."
192
+ },
193
+ {
194
+ "id": "axstack-arena-judge-astra",
195
+ "name": "Axstack arena judge Astra",
196
+ "provider": "codex",
197
+ "model": "gpt-6-astra",
198
+ "modeId": "full-access",
199
+ "thinkingOptionId": "xhigh",
200
+ "notes": "Arena judge (Astra seat). Read-only cross-judge for an arena-grade Align decision: receives the rubric and both candidates by label, scores each criterion, and recommends a base with rationale. Never authors a candidate, never cross-reads the other judge, never mutates. Disagreement between judges is surfaced by the driver, never averaged."
201
+ },
202
+ {
203
+ "id": "axstack-arena-judge-fable",
204
+ "name": "Axstack arena judge Fable",
205
+ "provider": "claude",
206
+ "model": "claude-fable-5-1",
207
+ "modeId": "bypassPermissions",
208
+ "thinkingOptionId": "xhigh",
209
+ "notes": "Arena judge (Fable seat). Read-only cross-judge for an arena-grade Align decision: receives the rubric and both candidates by label, scores each criterion, and recommends a base with rationale. Never authors a candidate, never cross-reads the other judge, never mutates. Disagreement between judges is surfaced by the driver, never averaged."
192
210
  }
193
211
  ]
194
212
  }
@@ -2,7 +2,7 @@
2
2
 
3
3
  Read this when the current session is the native Orca PR **driver** or its
4
4
  **watchdog**. The approved contract is
5
- `docs/specs/pr-automations.md` revision 4; this reference restates the parts an
5
+ `docs/specs/pr-automations.md` revision 5; this reference restates the parts an
6
6
  automation session must execute and does not widen them.
7
7
 
8
8
  Pair A/B is retired for this contract; its artefacts remain untouched.
@@ -127,32 +127,33 @@ The driver performs this order and exits:
127
127
  Claude Code trusts a folder per git toplevel and stops at its "Quick safety
128
128
  check" dialog otherwise, and the driver never answers that dialog for a
129
129
  worker. So workers run in a fixed pool: every allowlisted project has
130
- exactly two slot worktrees, `slot-1` and `slot-2`, its Orca child worktrees
131
- created once, parented to the project's primary worktree, and trusted once
130
+ exactly five slot worktrees, `slot-1` through `slot-5`, its Orca child
131
+ worktrees created once, parented to the project's primary worktree, and trusted once
132
132
  by the user through that dialog. The driver never creates or removes a
133
133
  worktree and never writes `~/.claude.json`. A slot is free when no live
134
134
  dispatch marker names it, it is not in `retained_slots[]`, no terminal is
135
135
  listed in it, and its tree is clean; no free slot in the project defers the
136
136
  PR to `deferred[]` like a budget, not a hold. Taking a slot fetches the head into the project clone,
137
137
  checks the slot out detached at the pinned head, and verifies HEAD equals
138
- it; the marker's worktree is the slot path. Trusting a folder activates the
139
- full project surface: its `.claude/settings.json` and the hooks it defines,
140
- its `.mcp.json` servers, marketplace plugin auto-install, and `CLAUDE.md`;
141
- a hostile branch's hooks or MCP servers would run the moment the folder
142
- opens. So the worker is launched in Claude Code's own isolation mode,
143
- `--safe-mode`, through `terminal create` and `worker-start --terminal`,
144
- since `worker-start` cannot pass argv. Safe mode is the binary's sanctioned
145
- "all customizations disabled" path: no `CLAUDE.md`, skills, plugins, hooks,
146
- MCP servers, custom commands or agents load from anywhere, project or user;
147
- built-in tools and authentication are untouched, and the brief loads the
148
- skill files it needs by path. The driver confirms readiness from the
149
- rendered frame — `wait.satisfied`, the prompt marker present, the dialog
150
- absent — before dispatching; a dialog on a slot means the slot is not
151
- trusted: the slot is named in a health line, the PR is deferred, nothing is
152
- dispatched. What remains live is the repository's files as data the worker
153
- reads and the commands the worker itself chooses to run, which it already
154
- runs today; nothing from the branch loads, executes, or is offered for
155
- invocation on its own.
138
+ it; the marker's worktree is the slot path. The worker is launched by Orca
139
+ itself — `worker-start --agent claude --model claude-opus-5 --effort medium`
140
+ on that slot — so the runtime owns the process: `worker-release` ends it and
141
+ `worker-show` proves it exited, which is what returns the slot to the pool.
142
+ Never pre-create the worker's terminal or hand a terminal handle to
143
+ `worker-start`: a reused handle is a resource Orca labels `external`, one it
144
+ can neither stop nor prove exited, so every such slot ends retained. After
145
+ `worker-start` run `worker-show` on the receipt's dispatch id and require
146
+ `projection.resource.state == owned`; anything else is a launch Orca does not
147
+ own: apply the runtime-refusal recovery rules under "Safety holds" (the
148
+ `worker-list` row's `nextAction` argv verbatim; `none` means inspect and
149
+ retain), append a `worker not owned` health line and a `retained_slots[]`
150
+ entry for the slot, defer the PR, and record no marker. Orca's per-agent default arguments supply
151
+ `--dangerously-skip-permissions`; the brief loads the skill files it needs by
152
+ path. If `worker-start` reports a failed stage or a visible hold (the "Quick
153
+ safety check" trust dialog) the slot is not trusted: name the slot in a health
154
+ line, defer the PR, dispatch nothing, never answer the dialog. Project
155
+ customizations load as they would for the user; the allowlist is defi-com
156
+ only and the user accepted that surface on 2026-09-18.
156
157
 
157
158
  Every selected PR receives one dispatch marker with task id, dispatch id,
158
159
  worktree, head, `started_at`, reservation (`verdict` or `repair`), and trigger:
@@ -0,0 +1,64 @@
1
+ # Blast radius
2
+
3
+ Find what a change breaks somewhere else before it ships, beyond the diff, and
4
+ prove the one fact it is safe because of by running real code instead of
5
+ writing it up. Loaded by `axstack-review` under angle 3 before any verdict and
6
+ by `axstack-explain` for a direct "what could this break" question.
7
+
8
+ Listing the callers is not the job; a symbol search does that in a second. The
9
+ job is the breakage the search does not show.
10
+
11
+ ## Evidence ladder
12
+
13
+ A writeup that sounds right is worthless: it reads as convincing whether or not
14
+ it is true. For each fact the change's safety depends on, get it as far down
15
+ this ladder as is cheap and say where it stopped:
16
+
17
+ 1. **Said so.** Worthless on its own.
18
+ 2. **Pointed at the line.** A real `file:line`, or the library's own source at
19
+ the pinned version.
20
+ 3. **Walked the failure.** The bad case was traced step by step and does not
21
+ reach.
22
+ 4. **Ran it.** A script or test that calls the real code and fails loud if the
23
+ claim is wrong.
24
+ 5. **Reproduced it in the running app.**
25
+
26
+ A safety fact below step 4 is written as **unproven**, never as settled. Step 4
27
+ is usually one small script that imports what the app ships and calls the exact
28
+ function in question. "Not found" is an answer for the inspected scope only.
29
+
30
+ ## Steps
31
+
32
+ 1. **Read the change.** The diff, the symbols it adds, changes, and deletes,
33
+ and what it now does differently, including the part the diff does not
34
+ spell out. Pull the PR body and commit messages for the stated intent.
35
+ 2. **Find the one fact it is safe because of.** Most changes that look risky
36
+ are safe because of a single fact, such as "this call only drops
37
+ already-dead cache entries". Spend the time here, not on a long list of
38
+ maybes. If that fact holds, most risky cases clear at once.
39
+ 3. **Look where the search stops.** Library source at the pinned version and
40
+ any local patch; when things run (microtasks, unmount and teardown,
41
+ reactive frameworks); what a symbol search misses: the JSON an API returns,
42
+ a database column, a wire or on-chain format, another language reading the
43
+ same bytes, a feature flag, code three hops downstream.
44
+ 4. **Be honest about each risk.** Give it a real chance of happening and a
45
+ real cost if it does. Keep the confirmed risks; list the checked and cleared
46
+ ones separately. Cite a real `file:line`; never invent a caller or an API.
47
+ 5. **Prove the one fact.** Write the script or test, run it against the real
48
+ code, and paste what happened. If it cannot be proven cheaply, mark it
49
+ unproven; do not overstate.
50
+
51
+ ## What to hand back
52
+
53
+ - **What it does.** What changed, including the part that is not obvious.
54
+ - **The one fact it is safe because of.** State it, the ladder step reached,
55
+ and the proof; or `unproven`.
56
+ - **Risks.** Only the real ones, each with how it breaks, the `file:line`, how
57
+ likely and how bad, and how to check.
58
+ - **Cleared.** What was checked and why it is fine.
59
+ - **Before merge.** The cheapest test or repro that catches the real bug,
60
+ including the script written for step 5.
61
+
62
+ Strip anything private before the writeup leaves the run. In review, an
63
+ unproven safety fact for a consequential change is a finding with its
64
+ consequence stated, not a footnote.
@@ -49,7 +49,10 @@ profile. Record the driver's provider and model in the run record.
49
49
 
50
50
  For Align and Spec, the driver forms an independent assessment first, then
51
51
  consults `axstack-advisor-astra` and `axstack-advisor-fable` independently with
52
- the same bounded evidence and question. The driver synthesizes disagreements,
52
+ the same bounded evidence and question. The one exception is an arena-grade
53
+ Align question: there the driver frames the brief and rubric, the advisers
54
+ author candidates, and the driver assesses only after the candidates and judge
55
+ verdicts return. The driver synthesizes disagreements,
53
56
  owns the decision, and the user still approves the spec. Reuse each valid
54
57
  unchanged receipt; changed evidence, scope, or question requires a fresh
55
58
  receipt. If either adviser is unavailable, Align and Spec hold without model or
@@ -26,7 +26,7 @@ Read `roles.json` relative to the actually loaded `axstack` skill. The installed
26
26
  shape is `{ "version": 1, "preset": "<name>", "roles": [...] }`. Bundled
27
27
  profiles are setup inputs shaped as
28
28
  `{ "version": 1, "roles": [...] }`. A new run records the selected preset and
29
- all 21 role rows once. An active run keeps the exact snapshot until the user
29
+ all 23 role rows once. An active run keeps the exact snapshot until the user
30
30
  explicitly changes it.
31
31
 
32
32
  Select the requested role by stable ID. A missing or null model holds only that role;
@@ -11,7 +11,7 @@ only with exactly one unambiguous preset; missing or contradictory sources are
11
11
  a setup gap: hold. Never infer from live profiles or `list_profiles`, harness,
12
12
  tools, credentials, quota, subscription, or default to `mixed`.
13
13
 
14
- At run start, capture one **routing snapshot**: the complete map of all 21 role
14
+ At run start, capture one **routing snapshot**: the complete map of all 23 role
15
15
  IDs with provider/model/mode/effort, absent or unconfigured roles recorded
16
16
  explicitly, and no invented provider default. An absent or unconfigured role
17
17
  holds only that role's work, not the run. A role installed or changed later
@@ -38,8 +38,10 @@ Role IDs:
38
38
  | `mixed` | Claude / Opus (`claude/claude-opus-5`) | `axstack-reviewer-primary` (`codex/gpt-5.6-sol` medium) |
39
39
  | `codex-only` | Codex / Sol (`codex/gpt-5.6-sol`) | `axstack-reviewer-secondary` (`codex/gpt-5.6-terra` xhigh) |
40
40
  | `claude-only` | Claude / Opus (`claude/claude-opus-5`) | `axstack-reviewer-secondary` (`claude/claude-sonnet-5` xhigh) |
41
- - `axstack-advisor-astra` and `axstack-advisor-fable` advise independently;
42
- `axstack-auditor` audits; `axstack-checker` reports discrepancies.
41
+ - `axstack-advisor-astra` and `axstack-advisor-fable` advise independently
42
+ and author align arena candidates; `axstack-arena-judge-astra` and
43
+ `axstack-arena-judge-fable` judge them. `axstack-auditor` audits;
44
+ `axstack-checker` reports discrepancies.
43
45
  - `axstack-explainer` authors explanations; `axstack-explainer-review`
44
46
  reviews them. `axstack-monitor` observes only; `axstack-watchdog` sends
45
47
  only gate-authorized health escalations.
@@ -47,10 +49,9 @@ Role IDs:
47
49
 
48
50
  Provenance is matched on provider/model ID; record effort but never use it to
49
51
  create a mapping. Provenance absent from the preset's table row is
50
- unsupported and `INCOMPLETE` (including its secondary reviewer model, Astra,
51
- Luna, or Fable); report the exact gap and ask the user. Never derive a reverse
52
- pairing from slot position, driver, owner, or provider. Author and owner never
53
- review their own work.
52
+ unsupported and `INCOMPLETE`; report the exact gap and ask the user. Never
53
+ derive a reverse pairing from slot position, driver, owner, or provider.
54
+ Author and owner never review their own work.
54
55
 
55
56
  ## Direct routes (no spec ceremony)
56
57
 
@@ -58,37 +59,37 @@ review their own work.
58
59
  and code, return a cited note with limitations. Fan out only distinct
59
60
  questions.
60
61
  - Understanding a system, change, or implementation gap -> `axstack-explain`:
61
- show current/intended behavior, evidence dimensions, and bounded gaps; use
62
- project docs and verify rendered behavior when applicable. Stale axstack-docs
63
- is superseded. Publication needs separate authority.
62
+ current/intended behavior, evidence dimensions, and bounded gaps from project
63
+ docs and rendered behavior. "What could this break" follows
64
+ [Blast radius](blast-radius.md). Publication needs separate authority.
64
65
  - A bug, failing test, regression, or wrong behavior, red loop wanted ->
65
66
  `axstack-debug`: diagnose, escalate via adviser-directed investigators, hand
66
67
  off a classified repair (explain: how; debug: what's wrong).
67
68
  - Codebase-quality or refactor discovery -> `axstack-improve`: inspect bounded
68
69
  scope, rank evidenced candidates, report only; no spec, tickets, or source
69
- edits. Selected changes return via preparation or execution.
70
+ edits.
70
71
  - Preparation completion, watch expiry, resume, or reconciliation -> the
71
72
  [lifecycle](lifecycle.md#native-handoff-and-resume): reconcile the run
72
73
  record, keep its owner, launch no native handoff.
73
74
  - Explicit user-requested ownership transfer -> the same lifecycle section.
74
75
  Load the [Orca runtime boundary](orca-runtime.md), follow the runtime-owned
75
76
  handoff guide, and require explicit recipient acceptance before ownership
76
- changes. Missing capability is a setup gap, never license to invent one.
77
+ changes. Missing capability is a setup gap; never invent one.
77
78
  - Colleague PR review -> `axstack-review`, peer mode.
78
79
  - Own PR maintenance or monitoring -> `axstack-review` in authored mode,
79
80
  `axstack-watch` for adoption.
80
81
 
81
82
  Research, explanation, improvement discovery, debugging, handoff, peer review,
82
- and adopted maintenance need no alignment, approved spec, or ticket map;
83
- authority and intent boundaries still apply.
83
+ and adopted maintenance need no alignment, spec, or ticket map; authority and
84
+ intent boundaries still apply.
84
85
 
85
86
  ## Proportional scope identity
86
87
 
87
88
  Classify new work as substantial, small, or unclear; record it with brief
88
89
  reason in the run record, or in the brief for tiny direct work.
89
90
 
90
- - **Substantial:** substantial features, multi-PR work, or stacked work. A
91
- bounded small feature is not substantial merely because it is labelled one.
91
+ - **Substantial:** substantial features, multi-PR work, or stacked work; a
92
+ bounded small feature is not substantial because it is labelled one.
92
93
  Require an approved spec plus a ticket map tied to that exact spec
93
94
  revision, with acceptance checks and dependencies in the explicitly selected
94
95
  Markdown or Linear store. Prepare via `axstack-align` -> `axstack-spec`
@@ -107,8 +108,8 @@ reason in the run record, or in the brief for tiny direct work.
107
108
  Reassess size when growth adds an additional PR, a new execution dependency
108
109
  that materially expands scope, an unsettled material design question, or a
109
110
  security/infrastructure boundary crossing. An ordinary
110
- test-then-code sequence is not multi-task growth; a minor file dependency does
111
- not alone need a formal spec. Hold affected unsafe work while reassessing.
111
+ test-then-code sequence is not multi-task growth; a minor file dependency is
112
+ not alone a formal spec trigger. Hold affected unsafe work while reassessing.
112
113
 
113
114
  ## Lifecycle routes (mode-specific scope identity required)
114
115
 
@@ -125,5 +126,5 @@ not alone need a formal spec. Hold affected unsafe work while reassessing.
125
126
  author provenance — then use `axstack-review` and `axstack-watch` without
126
127
  repeated approval or new spec ceremony. Never infer the author from the
127
128
  orchestrator or assume an imported own PR's author.
128
- - Direct later phase: start there and pass that phase's identity check. Entry
129
- never admits work a deeper phase would reject.
129
+ - Direct later phase: start there and pass that phase's identity check; entry
130
+ never admits work a deeper phase rejects.
@@ -74,6 +74,23 @@ Reconcile named sessions, revisions, PR state, watches, and deliveries before
74
74
  creating or redelivering anything. Touch only this run; no global sweep, new
75
75
  runtime database, or scheduler follows from the record.
76
76
 
77
+ ## Decision trail and learnings
78
+
79
+ Two fields make the record reviewable by a human who stepped away and usable
80
+ by the next owner:
81
+
82
+ - `Decisions` holds one row per consequential choice: when, what was chosen,
83
+ why in plain words, the evidence pointer that proves it (SHA, PR, receipt,
84
+ `file:line`, or artifact path, never a paragraph), and the result
85
+ (`tests green`, `reverted`, `held`, `open`). A choice that a reviewer could
86
+ not reconstruct from Git or PR state belongs here; routine mechanics do not.
87
+ An arena synthesis note (base, grafts and their source candidate,
88
+ rejections, judge verdicts) is recorded as `Decisions` rows.
89
+ - `Learnings` holds what the next owner needs that the code and history do not
90
+ show: root causes, gotchas, patterns that held or failed, and where the
91
+ evidence lives. Read it on resume before reconciling; append, never rewrite,
92
+ and keep entries as short as the evidence pointer allows.
93
+
77
94
  ## Privacy
78
95
 
79
96
  Record concise IDs, SHAs, URLs, status, timestamps, next actions, and evidence
@@ -98,8 +115,13 @@ IDs: <repo/project + workspace/agent receipt pointers>
98
115
  Evidence: <check/review/submission/audit receipt pointers>
99
116
  Pending: <launch/acceptance/external receipts + timer execution heartbeat actual ID + handshake + deadline>
100
117
  Unresolved: <decision -> next owner + next action>
118
+ Learnings: <root cause, gotcha, or pattern -> evidence ref>
101
119
  Resume: <commands or evidence refs bound to exact revisions>
102
120
 
121
+ | When | Decision | Why | Evidence | Result |
122
+ | --- | --- | --- | --- | --- |
123
+ | <UTC timestamp> | <what was chosen> | <plain reason> | <SHA/PR/receipt/file:line> | <tests green/reverted/held/open> |
124
+
103
125
  | Task | Dependencies | Owner | State | Revision evidence | Next action |
104
126
  | --- | --- | --- | --- | --- | --- |
105
127
  | <task> | <task IDs or none> | <role + session ID + worktree, or receipt ref> | <pending/in progress/complete/blocked> | <SHA + check/receipt refs> | <action + owner> |
@@ -41,7 +41,9 @@ This preserves the required contracts -> lifecycle -> audit load edge.
41
41
  The current chat remains the driver under
42
42
  [Standing contracts](../axstack/references/contracts.md). For each new
43
43
  user round, the driver independently drafts the prioritized frontier and
44
- recommendations. Then consult `axstack-advisor-astra` and
44
+ recommendations, except for an arena-grade question (below), where the driver
45
+ writes the brief and rubric but drafts no recommendation until the candidates
46
+ and judge verdicts return, so nothing anchors them. Then consult `axstack-advisor-astra` and
45
47
  `axstack-advisor-fable` independently, without cross-reading, using the same
46
48
  bounded evidence and question. Each adviser challenges assumptions, edges,
47
49
  omissions, and alternatives; the driver synthesizes disagreements and accepts
@@ -56,6 +58,48 @@ unchanged receipts. Record compact adviser evidence, the driver's assessment,
56
58
  and user-resolved choices for `axstack-spec`. If either adviser is unavailable,
57
59
  hold Align; safe fact work may continue without substitution.
58
60
 
61
+ ## Arena for hard-to-reverse design choices
62
+
63
+ Critique of one draft anchors every reader to that draft's shape. When a
64
+ question is arena-grade, the same test as for an ADR (a meaningful,
65
+ hard-to-reverse, non-obvious trade-off: architecture, module boundaries, data
66
+ model, migration strategy), replace the critique round for that question with
67
+ one arena round. Small or routine questions never enter the arena.
68
+
69
+ 1. **Frame.** The driver writes the brief (the artifact, its constraints, the
70
+ settled decisions it must respect) and three to six gradeable rubric
71
+ criteria. Candidates receive only the brief; the rubric is for judging.
72
+ 2. **Fan out.** `axstack-advisor-astra` and `axstack-advisor-fable` each
73
+ independently produce one candidate design plus a short rationale naming
74
+ the alternatives considered and rejected, from the same bounded evidence and
75
+ question, without cross-reading. The driver authors no candidate.
76
+ 3. **Cross-judge.** After both candidates are complete, `axstack-arena-judge-astra`
77
+ and `axstack-arena-judge-fable` each independently score every candidate
78
+ per criterion from the rubric and candidates by label, and recommend a base
79
+ with a reason. Judges never author, never cross-read each other.
80
+ 4. **Pick.** The driver reads every candidate end to end and scores per
81
+ criterion, not on holistic feel, then compares with both judges. Agreement
82
+ confirms the base. Disagreement between judges or with the driver means one
83
+ reading is biased or the rubric was ambiguous: re-read both rationales and
84
+ decide with a stated reason; never average verdicts or fabricate consensus.
85
+ 5. **Graft.** Walk the losing candidate once more for the one or two ideas
86
+ worth porting and fold them into the base by hand so the result stays
87
+ coherent under one mental model. Convergence on the same shape is a strong
88
+ agreement signal: adopt the consensus shape, no graft. Wide divergence
89
+ means the frame was under-specified: reframe and rerun once, never
90
+ average.
91
+ 6. **Present.** The synthesized design is the recommendation in the next
92
+ `Qn`, with its trade-off, judge verdicts, and what was grafted or rejected.
93
+ The user still decides; spec approval remains the one human checkpoint.
94
+
95
+ Record the synthesis note (base, grafts and their source candidate, rejections,
96
+ dropouts, both judge verdicts) as `Decisions` rows in the
97
+ [run record](../axstack/references/run-record.md). Load
98
+ [Orca runtime](../axstack/references/orca-runtime.md) immediately before the
99
+ first candidate or judge dispatch. If either adviser or judge seat is
100
+ unavailable, hold that question without substitution; unaffected fact work
101
+ and questions continue.
102
+
59
103
  ## Bound the interview
60
104
 
61
105
  Twenty cumulative presented questions is the normal ceiling, not a target.
@@ -33,6 +33,18 @@ immediately before an actual profile dispatch.
33
33
  4. For every gap, cite its inspected scope and applicable revision, stable
34
34
  source identity, or content hash. “Not found” never means app-wide missing
35
35
  without app-wide evidence; anything outside the inspected scope is unknown.
36
+ 5. For "what could this break" or "blast radius of X", follow
37
+ [Blast radius](../axstack/references/blast-radius.md): find the breakage
38
+ beyond the diff and prove the one safety fact by running real code, or
39
+ mark it unproven.
40
+ 6. For "show me your work" on existing work (a run, PR, branch, or change),
41
+ reconstruct the decision trail rather than re-describing the diff: read the
42
+ run record's `Decisions` and `Learnings`
43
+ ([run record](../axstack/references/run-record.md)), then the commits, PR
44
+ body, review receipts, and ADRs. Present each consequential choice as what
45
+ was chosen, why, the evidence pointer, and its result; separate what the
46
+ record proves from what is inferred, and list choices with no recorded
47
+ reason as open rather than inventing one.
36
48
 
37
49
  ## 2. Choose proportional output
38
50
 
@@ -134,7 +134,9 @@ session is the owner for every PR it handles.
134
134
  in authored mode the one reviewer covers all six angles:
135
135
  1. Security and trust boundaries.
136
136
  2. Correctness, failures, and edge cases.
137
- 3. Integration and regressions.
137
+ 3. Integration and regressions: load [Blast radius](../axstack/references/blast-radius.md),
138
+ find what the change breaks beyond the diff, and grade the one fact it
139
+ is safe because of on the evidence ladder; below "ran it" is unproven.
138
140
  4. Requirements, acceptance, and user behavior.
139
141
  5. Architecture and solution design, including SOLID and credible simpler
140
142
  alternatives.
@@ -233,6 +235,7 @@ Verdict: <APPROVE | REQUEST_CHANGES | INCOMPLETE>
233
235
  Coverage: <angles + acceptance + executable evidence checked>
234
236
  Limitations: <unverified boundaries + why>
235
237
  Findings: <evidence + consequence each>
238
+ Safety fact: <the one fact the change is safe because of> — <ladder step + proof | unproven>
236
239
  Escalate to user: <yes | no> — <criterion> — <reason>
237
240
  ```
238
241
 
package/src/roles.js CHANGED
@@ -64,8 +64,8 @@ export function assessRoleReadiness(roles, preset) {
64
64
  const gaps = [];
65
65
  const isIntentionalAbsence = (role) => role.model === null && (
66
66
  (preset === 'mixed' && role.id === 'axstack-checker') ||
67
- (preset === 'codex-only' && role.id === 'axstack-advisor-fable') ||
68
- (preset === 'claude-only' && role.id === 'axstack-advisor-astra')
67
+ (preset === 'codex-only' && ['axstack-advisor-fable', 'axstack-arena-judge-fable'].includes(role.id)) ||
68
+ (preset === 'claude-only' && ['axstack-advisor-astra', 'axstack-arena-judge-astra'].includes(role.id))
69
69
  );
70
70
  for (const role of roles) {
71
71
  if (!bounds.has(role.provider)) {