axstack 0.13.1 → 0.14.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +1 -1
- package/docs/installation.md +9 -8
- package/docs/workflows.md +8 -2
- package/package.json +1 -1
- package/profiles/presets/claude-only.json +18 -0
- package/profiles/presets/codex-only.json +18 -0
- package/profiles/presets/mixed.json +18 -0
- package/skills/axstack/references/automations.md +22 -21
- package/skills/axstack/references/blast-radius.md +64 -0
- package/skills/axstack/references/contracts.md +4 -1
- package/skills/axstack/references/orca-runtime.md +1 -1
- package/skills/axstack/references/routing.md +21 -20
- package/skills/axstack/references/run-record.md +22 -0
- package/skills/axstack-align/SKILL.md +45 -1
- package/skills/axstack-explain/SKILL.md +12 -0
- package/skills/axstack-review/SKILL.md +4 -1
- package/src/roles.js +2 -2
package/README.md
CHANGED
|
@@ -77,7 +77,7 @@ targets are reported without normal-path adoption. Install exits nonzero when
|
|
|
77
77
|
an instruction conflict is preserved, while clean and idempotent installs exit
|
|
78
78
|
successfully.
|
|
79
79
|
|
|
80
|
-
The public bundle preserves three canonical
|
|
80
|
+
The public bundle preserves three canonical 23-role inputs:
|
|
81
81
|
[mixed](profiles/presets/mixed.json),
|
|
82
82
|
[codex-only](profiles/presets/codex-only.json), and
|
|
83
83
|
[claude-only](profiles/presets/claude-only.json). Each is exactly
|
package/docs/installation.md
CHANGED
|
@@ -51,7 +51,7 @@ profiles/presets/codex-only.json
|
|
|
51
51
|
profiles/presets/claude-only.json
|
|
52
52
|
```
|
|
53
53
|
|
|
54
|
-
Each has exactly `{ "version": 1, "roles": [...] }` with the same
|
|
54
|
+
Each has exactly `{ "version": 1, "roles": [...] }` with the same 23 stable
|
|
55
55
|
role IDs. Installation writes `<skills-dir>/axstack/roles.json` as
|
|
56
56
|
`{ "version": 1, "preset": "<selected preset>", "roles": [...] }` and records
|
|
57
57
|
its ownership hash like every other installed skill asset. There is no second
|
|
@@ -118,9 +118,9 @@ The complete bundle is validated before writes:
|
|
|
118
118
|
- each preset is a real JSON file with version 1, a non-empty `roles` array,
|
|
119
119
|
the filename's selected identity supplied by the caller, and the same role-ID
|
|
120
120
|
set as its peers;
|
|
121
|
-
- every role has valid preserved fields, while the mixed checker and
|
|
122
|
-
unavailable adviser
|
|
123
|
-
`model: null`;
|
|
121
|
+
- every role has valid preserved fields, while the mixed checker and, in each
|
|
122
|
+
single-provider preset, the unavailable adviser and its matching arena judge
|
|
123
|
+
seat explicitly permit `model: null`;
|
|
124
124
|
- obsolete runtime configuration flags fail before mutation with migration
|
|
125
125
|
guidance.
|
|
126
126
|
|
|
@@ -148,15 +148,16 @@ to rewrite them.
|
|
|
148
148
|
## Role behavior after installation
|
|
149
149
|
|
|
150
150
|
The runtime reads `roles.json` relative to the actually loaded `axstack` skill.
|
|
151
|
-
A new run records the selected preset plus all
|
|
151
|
+
A new run records the selected preset plus all 23 role rows. An active run keeps
|
|
152
152
|
that snapshot after a later preset install unless the user explicitly changes
|
|
153
153
|
it and accepts the resulting evidence invalidation.
|
|
154
154
|
|
|
155
155
|
The mixed checker has `model: null`; checker dispatch is held and never inherits
|
|
156
156
|
a provider default. The single-provider presets configure the checker. Their
|
|
157
|
-
unavailable adviser
|
|
158
|
-
|
|
159
|
-
and
|
|
157
|
+
unavailable adviser and its matching arena judge seat remain explicit
|
|
158
|
+
same-provider `model: null` roles, which do not make installation unready;
|
|
159
|
+
Align and Spec still hold until both Astra and Fable can return independent
|
|
160
|
+
receipts, and an arena-grade Align question holds until both judge seats can. The current chat drives on whatever
|
|
160
161
|
model runs it; no preset carries a driver role. Every other missing, invalid, unsupported, or unavailable role value holds only
|
|
161
162
|
the affected work. There is no model substitution, subscription inference, or
|
|
162
163
|
quota routing.
|
package/docs/workflows.md
CHANGED
|
@@ -40,7 +40,7 @@ only affected work.
|
|
|
40
40
|
|
|
41
41
|
Installation requires one explicit canonical preset. The three bundle files
|
|
42
42
|
under `profiles/presets/` each contain exactly
|
|
43
|
-
`{ "version": 1, "roles": [...] }` and the same
|
|
43
|
+
`{ "version": 1, "roles": [...] }` and the same 23 stable IDs.
|
|
44
44
|
|
|
45
45
|
The current chat drives on whatever model runs it; no preset carries a driver
|
|
46
46
|
role.
|
|
@@ -108,7 +108,13 @@ session and evidence remain valid.
|
|
|
108
108
|
|
|
109
109
|
- `axstack-align` maps facts and dependencies, asks prioritized questions, and
|
|
110
110
|
consults Astra and Fable independently with the same bounded evidence and
|
|
111
|
-
question. It synthesizes disagreements and reuses unchanged receipts.
|
|
111
|
+
question. It synthesizes disagreements and reuses unchanged receipts. For a
|
|
112
|
+
hard-to-reverse design choice it runs one arena round instead: Astra and
|
|
113
|
+
Fable each author a candidate, `axstack-arena-judge-astra` and
|
|
114
|
+
`axstack-arena-judge-fable` score both against the driver's rubric, and the
|
|
115
|
+
driver picks a base, grafts the losers' strong ideas, and presents the
|
|
116
|
+
synthesis as the recommendation; the note lands as `Decisions` rows in the
|
|
117
|
+
run record.
|
|
112
118
|
- `axstack-spec` writes observable acceptance, exclusions, decisions, and one
|
|
113
119
|
user-approved revision baseline.
|
|
114
120
|
- `axstack-tickets` maps user-visible capabilities to dependency-aware internal
|
package/package.json
CHANGED
|
@@ -189,6 +189,24 @@
|
|
|
189
189
|
"modeId": "bypassPermissions",
|
|
190
190
|
"thinkingOptionId": "xhigh",
|
|
191
191
|
"notes": "Debug investigator seat 4. Dispatched only by axstack-debug at L1 with the shared evidence packet and one distinct brief; never reads another investigator's output. Works in its own disposable worktree at the pinned revision plus the recorded dirty patch; may instrument there for probes; never commits, pushes, publishes, or creates children. Returns one receipt per brief. Independence comes from brief isolation, not model diversity. This preset repeats claude-sonnet-5 at xhigh effort because it has fewer model families; independence comes from brief isolation."
|
|
192
|
+
},
|
|
193
|
+
{
|
|
194
|
+
"id": "axstack-arena-judge-astra",
|
|
195
|
+
"name": "Axstack arena judge Astra",
|
|
196
|
+
"provider": "claude",
|
|
197
|
+
"model": null,
|
|
198
|
+
"modeId": "bypassPermissions",
|
|
199
|
+
"thinkingOptionId": "xhigh",
|
|
200
|
+
"notes": "Required Arena judge (Astra seat). Read-only cross-judge for an arena-grade Align decision: receives the rubric and both candidates by label, scores each criterion, and recommends a base with rationale. Never authors a candidate, never cross-reads the other judge, never mutates. Disagreement between judges is surfaced by the driver, never averaged. Intentionally absent in the claude-only preset; the arena-grade decision holds without substitution."
|
|
201
|
+
},
|
|
202
|
+
{
|
|
203
|
+
"id": "axstack-arena-judge-fable",
|
|
204
|
+
"name": "Axstack arena judge Fable",
|
|
205
|
+
"provider": "claude",
|
|
206
|
+
"model": "claude-fable-5-1",
|
|
207
|
+
"modeId": "bypassPermissions",
|
|
208
|
+
"thinkingOptionId": "xhigh",
|
|
209
|
+
"notes": "Arena judge (Fable seat). Read-only cross-judge for an arena-grade Align decision: receives the rubric and both candidates by label, scores each criterion, and recommends a base with rationale. Never authors a candidate, never cross-reads the other judge, never mutates. Disagreement between judges is surfaced by the driver, never averaged."
|
|
192
210
|
}
|
|
193
211
|
]
|
|
194
212
|
}
|
|
@@ -189,6 +189,24 @@
|
|
|
189
189
|
"modeId": "full-access",
|
|
190
190
|
"thinkingOptionId": "xhigh",
|
|
191
191
|
"notes": "Debug investigator seat 4. Dispatched only by axstack-debug at L1 with the shared evidence packet and one distinct brief; never reads another investigator's output. Works in its own disposable worktree at the pinned revision plus the recorded dirty patch; may instrument there for probes; never commits, pushes, publishes, or creates children. Returns one receipt per brief. Independence comes from brief isolation, not model diversity. This preset repeats gpt-5.6-terra at xhigh effort because it has fewer model families."
|
|
192
|
+
},
|
|
193
|
+
{
|
|
194
|
+
"id": "axstack-arena-judge-astra",
|
|
195
|
+
"name": "Axstack arena judge Astra",
|
|
196
|
+
"provider": "codex",
|
|
197
|
+
"model": "gpt-6-astra",
|
|
198
|
+
"modeId": "full-access",
|
|
199
|
+
"thinkingOptionId": "xhigh",
|
|
200
|
+
"notes": "Arena judge (Astra seat). Read-only cross-judge for an arena-grade Align decision: receives the rubric and both candidates by label, scores each criterion, and recommends a base with rationale. Never authors a candidate, never cross-reads the other judge, never mutates. Disagreement between judges is surfaced by the driver, never averaged."
|
|
201
|
+
},
|
|
202
|
+
{
|
|
203
|
+
"id": "axstack-arena-judge-fable",
|
|
204
|
+
"name": "Axstack arena judge Fable",
|
|
205
|
+
"provider": "codex",
|
|
206
|
+
"model": null,
|
|
207
|
+
"modeId": "full-access",
|
|
208
|
+
"thinkingOptionId": "xhigh",
|
|
209
|
+
"notes": "Required Arena judge (Fable seat). Read-only cross-judge for an arena-grade Align decision: receives the rubric and both candidates by label, scores each criterion, and recommends a base with rationale. Never authors a candidate, never cross-reads the other judge, never mutates. Disagreement between judges is surfaced by the driver, never averaged. Intentionally absent in the codex-only preset; the arena-grade decision holds without substitution."
|
|
192
210
|
}
|
|
193
211
|
]
|
|
194
212
|
}
|
|
@@ -189,6 +189,24 @@
|
|
|
189
189
|
"modeId": "full-access",
|
|
190
190
|
"thinkingOptionId": "low",
|
|
191
191
|
"notes": "Debug investigator seat 4. Dispatched only by axstack-debug at L1 with the shared evidence packet and one distinct brief; never reads another investigator's output. Works in its own disposable worktree at the pinned revision plus the recorded dirty patch; may instrument there for probes; never commits, pushes, publishes, or creates children. Returns one receipt per brief. Independence comes from brief isolation, not model diversity."
|
|
192
|
+
},
|
|
193
|
+
{
|
|
194
|
+
"id": "axstack-arena-judge-astra",
|
|
195
|
+
"name": "Axstack arena judge Astra",
|
|
196
|
+
"provider": "codex",
|
|
197
|
+
"model": "gpt-6-astra",
|
|
198
|
+
"modeId": "full-access",
|
|
199
|
+
"thinkingOptionId": "xhigh",
|
|
200
|
+
"notes": "Arena judge (Astra seat). Read-only cross-judge for an arena-grade Align decision: receives the rubric and both candidates by label, scores each criterion, and recommends a base with rationale. Never authors a candidate, never cross-reads the other judge, never mutates. Disagreement between judges is surfaced by the driver, never averaged."
|
|
201
|
+
},
|
|
202
|
+
{
|
|
203
|
+
"id": "axstack-arena-judge-fable",
|
|
204
|
+
"name": "Axstack arena judge Fable",
|
|
205
|
+
"provider": "claude",
|
|
206
|
+
"model": "claude-fable-5-1",
|
|
207
|
+
"modeId": "bypassPermissions",
|
|
208
|
+
"thinkingOptionId": "xhigh",
|
|
209
|
+
"notes": "Arena judge (Fable seat). Read-only cross-judge for an arena-grade Align decision: receives the rubric and both candidates by label, scores each criterion, and recommends a base with rationale. Never authors a candidate, never cross-reads the other judge, never mutates. Disagreement between judges is surfaced by the driver, never averaged."
|
|
192
210
|
}
|
|
193
211
|
]
|
|
194
212
|
}
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Read this when the current session is the native Orca PR **driver** or its
|
|
4
4
|
**watchdog**. The approved contract is
|
|
5
|
-
`docs/specs/pr-automations.md` revision
|
|
5
|
+
`docs/specs/pr-automations.md` revision 5; this reference restates the parts an
|
|
6
6
|
automation session must execute and does not widen them.
|
|
7
7
|
|
|
8
8
|
Pair A/B is retired for this contract; its artefacts remain untouched.
|
|
@@ -127,32 +127,33 @@ The driver performs this order and exits:
|
|
|
127
127
|
Claude Code trusts a folder per git toplevel and stops at its "Quick safety
|
|
128
128
|
check" dialog otherwise, and the driver never answers that dialog for a
|
|
129
129
|
worker. So workers run in a fixed pool: every allowlisted project has
|
|
130
|
-
exactly
|
|
131
|
-
created once, parented to the project's primary worktree, and trusted once
|
|
130
|
+
exactly five slot worktrees, `slot-1` through `slot-5`, its Orca child
|
|
131
|
+
worktrees created once, parented to the project's primary worktree, and trusted once
|
|
132
132
|
by the user through that dialog. The driver never creates or removes a
|
|
133
133
|
worktree and never writes `~/.claude.json`. A slot is free when no live
|
|
134
134
|
dispatch marker names it, it is not in `retained_slots[]`, no terminal is
|
|
135
135
|
listed in it, and its tree is clean; no free slot in the project defers the
|
|
136
136
|
PR to `deferred[]` like a budget, not a hold. Taking a slot fetches the head into the project clone,
|
|
137
137
|
checks the slot out detached at the pinned head, and verifies HEAD equals
|
|
138
|
-
it; the marker's worktree is the slot path.
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
152
|
-
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
|
|
138
|
+
it; the marker's worktree is the slot path. The worker is launched by Orca
|
|
139
|
+
itself — `worker-start --agent claude --model claude-opus-5 --effort medium`
|
|
140
|
+
on that slot — so the runtime owns the process: `worker-release` ends it and
|
|
141
|
+
`worker-show` proves it exited, which is what returns the slot to the pool.
|
|
142
|
+
Never pre-create the worker's terminal or hand a terminal handle to
|
|
143
|
+
`worker-start`: a reused handle is a resource Orca labels `external`, one it
|
|
144
|
+
can neither stop nor prove exited, so every such slot ends retained. After
|
|
145
|
+
`worker-start` run `worker-show` on the receipt's dispatch id and require
|
|
146
|
+
`projection.resource.state == owned`; anything else is a launch Orca does not
|
|
147
|
+
own: apply the runtime-refusal recovery rules under "Safety holds" (the
|
|
148
|
+
`worker-list` row's `nextAction` argv verbatim; `none` means inspect and
|
|
149
|
+
retain), append a `worker not owned` health line and a `retained_slots[]`
|
|
150
|
+
entry for the slot, defer the PR, and record no marker. Orca's per-agent default arguments supply
|
|
151
|
+
`--dangerously-skip-permissions`; the brief loads the skill files it needs by
|
|
152
|
+
path. If `worker-start` reports a failed stage or a visible hold (the "Quick
|
|
153
|
+
safety check" trust dialog) the slot is not trusted: name the slot in a health
|
|
154
|
+
line, defer the PR, dispatch nothing, never answer the dialog. Project
|
|
155
|
+
customizations load as they would for the user; the allowlist is defi-com
|
|
156
|
+
only and the user accepted that surface on 2026-09-18.
|
|
156
157
|
|
|
157
158
|
Every selected PR receives one dispatch marker with task id, dispatch id,
|
|
158
159
|
worktree, head, `started_at`, reservation (`verdict` or `repair`), and trigger:
|
|
@@ -0,0 +1,64 @@
|
|
|
1
|
+
# Blast radius
|
|
2
|
+
|
|
3
|
+
Find what a change breaks somewhere else before it ships, beyond the diff, and
|
|
4
|
+
prove the one fact it is safe because of by running real code instead of
|
|
5
|
+
writing it up. Loaded by `axstack-review` under angle 3 before any verdict and
|
|
6
|
+
by `axstack-explain` for a direct "what could this break" question.
|
|
7
|
+
|
|
8
|
+
Listing the callers is not the job; a symbol search does that in a second. The
|
|
9
|
+
job is the breakage the search does not show.
|
|
10
|
+
|
|
11
|
+
## Evidence ladder
|
|
12
|
+
|
|
13
|
+
A writeup that sounds right is worthless: it reads as convincing whether or not
|
|
14
|
+
it is true. For each fact the change's safety depends on, get it as far down
|
|
15
|
+
this ladder as is cheap and say where it stopped:
|
|
16
|
+
|
|
17
|
+
1. **Said so.** Worthless on its own.
|
|
18
|
+
2. **Pointed at the line.** A real `file:line`, or the library's own source at
|
|
19
|
+
the pinned version.
|
|
20
|
+
3. **Walked the failure.** The bad case was traced step by step and does not
|
|
21
|
+
reach.
|
|
22
|
+
4. **Ran it.** A script or test that calls the real code and fails loud if the
|
|
23
|
+
claim is wrong.
|
|
24
|
+
5. **Reproduced it in the running app.**
|
|
25
|
+
|
|
26
|
+
A safety fact below step 4 is written as **unproven**, never as settled. Step 4
|
|
27
|
+
is usually one small script that imports what the app ships and calls the exact
|
|
28
|
+
function in question. "Not found" is an answer for the inspected scope only.
|
|
29
|
+
|
|
30
|
+
## Steps
|
|
31
|
+
|
|
32
|
+
1. **Read the change.** The diff, the symbols it adds, changes, and deletes,
|
|
33
|
+
and what it now does differently, including the part the diff does not
|
|
34
|
+
spell out. Pull the PR body and commit messages for the stated intent.
|
|
35
|
+
2. **Find the one fact it is safe because of.** Most changes that look risky
|
|
36
|
+
are safe because of a single fact, such as "this call only drops
|
|
37
|
+
already-dead cache entries". Spend the time here, not on a long list of
|
|
38
|
+
maybes. If that fact holds, most risky cases clear at once.
|
|
39
|
+
3. **Look where the search stops.** Library source at the pinned version and
|
|
40
|
+
any local patch; when things run (microtasks, unmount and teardown,
|
|
41
|
+
reactive frameworks); what a symbol search misses: the JSON an API returns,
|
|
42
|
+
a database column, a wire or on-chain format, another language reading the
|
|
43
|
+
same bytes, a feature flag, code three hops downstream.
|
|
44
|
+
4. **Be honest about each risk.** Give it a real chance of happening and a
|
|
45
|
+
real cost if it does. Keep the confirmed risks; list the checked and cleared
|
|
46
|
+
ones separately. Cite a real `file:line`; never invent a caller or an API.
|
|
47
|
+
5. **Prove the one fact.** Write the script or test, run it against the real
|
|
48
|
+
code, and paste what happened. If it cannot be proven cheaply, mark it
|
|
49
|
+
unproven; do not overstate.
|
|
50
|
+
|
|
51
|
+
## What to hand back
|
|
52
|
+
|
|
53
|
+
- **What it does.** What changed, including the part that is not obvious.
|
|
54
|
+
- **The one fact it is safe because of.** State it, the ladder step reached,
|
|
55
|
+
and the proof; or `unproven`.
|
|
56
|
+
- **Risks.** Only the real ones, each with how it breaks, the `file:line`, how
|
|
57
|
+
likely and how bad, and how to check.
|
|
58
|
+
- **Cleared.** What was checked and why it is fine.
|
|
59
|
+
- **Before merge.** The cheapest test or repro that catches the real bug,
|
|
60
|
+
including the script written for step 5.
|
|
61
|
+
|
|
62
|
+
Strip anything private before the writeup leaves the run. In review, an
|
|
63
|
+
unproven safety fact for a consequential change is a finding with its
|
|
64
|
+
consequence stated, not a footnote.
|
|
@@ -49,7 +49,10 @@ profile. Record the driver's provider and model in the run record.
|
|
|
49
49
|
|
|
50
50
|
For Align and Spec, the driver forms an independent assessment first, then
|
|
51
51
|
consults `axstack-advisor-astra` and `axstack-advisor-fable` independently with
|
|
52
|
-
the same bounded evidence and question. The
|
|
52
|
+
the same bounded evidence and question. The one exception is an arena-grade
|
|
53
|
+
Align question: there the driver frames the brief and rubric, the advisers
|
|
54
|
+
author candidates, and the driver assesses only after the candidates and judge
|
|
55
|
+
verdicts return. The driver synthesizes disagreements,
|
|
53
56
|
owns the decision, and the user still approves the spec. Reuse each valid
|
|
54
57
|
unchanged receipt; changed evidence, scope, or question requires a fresh
|
|
55
58
|
receipt. If either adviser is unavailable, Align and Spec hold without model or
|
|
@@ -26,7 +26,7 @@ Read `roles.json` relative to the actually loaded `axstack` skill. The installed
|
|
|
26
26
|
shape is `{ "version": 1, "preset": "<name>", "roles": [...] }`. Bundled
|
|
27
27
|
profiles are setup inputs shaped as
|
|
28
28
|
`{ "version": 1, "roles": [...] }`. A new run records the selected preset and
|
|
29
|
-
all
|
|
29
|
+
all 23 role rows once. An active run keeps the exact snapshot until the user
|
|
30
30
|
explicitly changes it.
|
|
31
31
|
|
|
32
32
|
Select the requested role by stable ID. A missing or null model holds only that role;
|
|
@@ -11,7 +11,7 @@ only with exactly one unambiguous preset; missing or contradictory sources are
|
|
|
11
11
|
a setup gap: hold. Never infer from live profiles or `list_profiles`, harness,
|
|
12
12
|
tools, credentials, quota, subscription, or default to `mixed`.
|
|
13
13
|
|
|
14
|
-
At run start, capture one **routing snapshot**: the complete map of all
|
|
14
|
+
At run start, capture one **routing snapshot**: the complete map of all 23 role
|
|
15
15
|
IDs with provider/model/mode/effort, absent or unconfigured roles recorded
|
|
16
16
|
explicitly, and no invented provider default. An absent or unconfigured role
|
|
17
17
|
holds only that role's work, not the run. A role installed or changed later
|
|
@@ -38,8 +38,10 @@ Role IDs:
|
|
|
38
38
|
| `mixed` | Claude / Opus (`claude/claude-opus-5`) | `axstack-reviewer-primary` (`codex/gpt-5.6-sol` medium) |
|
|
39
39
|
| `codex-only` | Codex / Sol (`codex/gpt-5.6-sol`) | `axstack-reviewer-secondary` (`codex/gpt-5.6-terra` xhigh) |
|
|
40
40
|
| `claude-only` | Claude / Opus (`claude/claude-opus-5`) | `axstack-reviewer-secondary` (`claude/claude-sonnet-5` xhigh) |
|
|
41
|
-
- `axstack-advisor-astra` and `axstack-advisor-fable` advise independently
|
|
42
|
-
|
|
41
|
+
- `axstack-advisor-astra` and `axstack-advisor-fable` advise independently
|
|
42
|
+
and author align arena candidates; `axstack-arena-judge-astra` and
|
|
43
|
+
`axstack-arena-judge-fable` judge them. `axstack-auditor` audits;
|
|
44
|
+
`axstack-checker` reports discrepancies.
|
|
43
45
|
- `axstack-explainer` authors explanations; `axstack-explainer-review`
|
|
44
46
|
reviews them. `axstack-monitor` observes only; `axstack-watchdog` sends
|
|
45
47
|
only gate-authorized health escalations.
|
|
@@ -47,10 +49,9 @@ Role IDs:
|
|
|
47
49
|
|
|
48
50
|
Provenance is matched on provider/model ID; record effort but never use it to
|
|
49
51
|
create a mapping. Provenance absent from the preset's table row is
|
|
50
|
-
unsupported and `INCOMPLETE
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
review their own work.
|
|
52
|
+
unsupported and `INCOMPLETE`; report the exact gap and ask the user. Never
|
|
53
|
+
derive a reverse pairing from slot position, driver, owner, or provider.
|
|
54
|
+
Author and owner never review their own work.
|
|
54
55
|
|
|
55
56
|
## Direct routes (no spec ceremony)
|
|
56
57
|
|
|
@@ -58,37 +59,37 @@ review their own work.
|
|
|
58
59
|
and code, return a cited note with limitations. Fan out only distinct
|
|
59
60
|
questions.
|
|
60
61
|
- Understanding a system, change, or implementation gap -> `axstack-explain`:
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
62
|
+
current/intended behavior, evidence dimensions, and bounded gaps from project
|
|
63
|
+
docs and rendered behavior. "What could this break" follows
|
|
64
|
+
[Blast radius](blast-radius.md). Publication needs separate authority.
|
|
64
65
|
- A bug, failing test, regression, or wrong behavior, red loop wanted ->
|
|
65
66
|
`axstack-debug`: diagnose, escalate via adviser-directed investigators, hand
|
|
66
67
|
off a classified repair (explain: how; debug: what's wrong).
|
|
67
68
|
- Codebase-quality or refactor discovery -> `axstack-improve`: inspect bounded
|
|
68
69
|
scope, rank evidenced candidates, report only; no spec, tickets, or source
|
|
69
|
-
edits.
|
|
70
|
+
edits.
|
|
70
71
|
- Preparation completion, watch expiry, resume, or reconciliation -> the
|
|
71
72
|
[lifecycle](lifecycle.md#native-handoff-and-resume): reconcile the run
|
|
72
73
|
record, keep its owner, launch no native handoff.
|
|
73
74
|
- Explicit user-requested ownership transfer -> the same lifecycle section.
|
|
74
75
|
Load the [Orca runtime boundary](orca-runtime.md), follow the runtime-owned
|
|
75
76
|
handoff guide, and require explicit recipient acceptance before ownership
|
|
76
|
-
changes. Missing capability is a setup gap
|
|
77
|
+
changes. Missing capability is a setup gap; never invent one.
|
|
77
78
|
- Colleague PR review -> `axstack-review`, peer mode.
|
|
78
79
|
- Own PR maintenance or monitoring -> `axstack-review` in authored mode,
|
|
79
80
|
`axstack-watch` for adoption.
|
|
80
81
|
|
|
81
82
|
Research, explanation, improvement discovery, debugging, handoff, peer review,
|
|
82
|
-
and adopted maintenance need no alignment,
|
|
83
|
-
|
|
83
|
+
and adopted maintenance need no alignment, spec, or ticket map; authority and
|
|
84
|
+
intent boundaries still apply.
|
|
84
85
|
|
|
85
86
|
## Proportional scope identity
|
|
86
87
|
|
|
87
88
|
Classify new work as substantial, small, or unclear; record it with brief
|
|
88
89
|
reason in the run record, or in the brief for tiny direct work.
|
|
89
90
|
|
|
90
|
-
- **Substantial:** substantial features, multi-PR work, or stacked work
|
|
91
|
-
bounded small feature is not substantial
|
|
91
|
+
- **Substantial:** substantial features, multi-PR work, or stacked work; a
|
|
92
|
+
bounded small feature is not substantial because it is labelled one.
|
|
92
93
|
Require an approved spec plus a ticket map tied to that exact spec
|
|
93
94
|
revision, with acceptance checks and dependencies in the explicitly selected
|
|
94
95
|
Markdown or Linear store. Prepare via `axstack-align` -> `axstack-spec`
|
|
@@ -107,8 +108,8 @@ reason in the run record, or in the brief for tiny direct work.
|
|
|
107
108
|
Reassess size when growth adds an additional PR, a new execution dependency
|
|
108
109
|
that materially expands scope, an unsettled material design question, or a
|
|
109
110
|
security/infrastructure boundary crossing. An ordinary
|
|
110
|
-
test-then-code sequence is not multi-task growth; a minor file dependency
|
|
111
|
-
not alone
|
|
111
|
+
test-then-code sequence is not multi-task growth; a minor file dependency is
|
|
112
|
+
not alone a formal spec trigger. Hold affected unsafe work while reassessing.
|
|
112
113
|
|
|
113
114
|
## Lifecycle routes (mode-specific scope identity required)
|
|
114
115
|
|
|
@@ -125,5 +126,5 @@ not alone need a formal spec. Hold affected unsafe work while reassessing.
|
|
|
125
126
|
author provenance — then use `axstack-review` and `axstack-watch` without
|
|
126
127
|
repeated approval or new spec ceremony. Never infer the author from the
|
|
127
128
|
orchestrator or assume an imported own PR's author.
|
|
128
|
-
- Direct later phase: start there and pass that phase's identity check
|
|
129
|
-
never admits work a deeper phase
|
|
129
|
+
- Direct later phase: start there and pass that phase's identity check; entry
|
|
130
|
+
never admits work a deeper phase rejects.
|
|
@@ -74,6 +74,23 @@ Reconcile named sessions, revisions, PR state, watches, and deliveries before
|
|
|
74
74
|
creating or redelivering anything. Touch only this run; no global sweep, new
|
|
75
75
|
runtime database, or scheduler follows from the record.
|
|
76
76
|
|
|
77
|
+
## Decision trail and learnings
|
|
78
|
+
|
|
79
|
+
Two fields make the record reviewable by a human who stepped away and usable
|
|
80
|
+
by the next owner:
|
|
81
|
+
|
|
82
|
+
- `Decisions` holds one row per consequential choice: when, what was chosen,
|
|
83
|
+
why in plain words, the evidence pointer that proves it (SHA, PR, receipt,
|
|
84
|
+
`file:line`, or artifact path, never a paragraph), and the result
|
|
85
|
+
(`tests green`, `reverted`, `held`, `open`). A choice that a reviewer could
|
|
86
|
+
not reconstruct from Git or PR state belongs here; routine mechanics do not.
|
|
87
|
+
An arena synthesis note (base, grafts and their source candidate,
|
|
88
|
+
rejections, judge verdicts) is recorded as `Decisions` rows.
|
|
89
|
+
- `Learnings` holds what the next owner needs that the code and history do not
|
|
90
|
+
show: root causes, gotchas, patterns that held or failed, and where the
|
|
91
|
+
evidence lives. Read it on resume before reconciling; append, never rewrite,
|
|
92
|
+
and keep entries as short as the evidence pointer allows.
|
|
93
|
+
|
|
77
94
|
## Privacy
|
|
78
95
|
|
|
79
96
|
Record concise IDs, SHAs, URLs, status, timestamps, next actions, and evidence
|
|
@@ -98,8 +115,13 @@ IDs: <repo/project + workspace/agent receipt pointers>
|
|
|
98
115
|
Evidence: <check/review/submission/audit receipt pointers>
|
|
99
116
|
Pending: <launch/acceptance/external receipts + timer execution heartbeat actual ID + handshake + deadline>
|
|
100
117
|
Unresolved: <decision -> next owner + next action>
|
|
118
|
+
Learnings: <root cause, gotcha, or pattern -> evidence ref>
|
|
101
119
|
Resume: <commands or evidence refs bound to exact revisions>
|
|
102
120
|
|
|
121
|
+
| When | Decision | Why | Evidence | Result |
|
|
122
|
+
| --- | --- | --- | --- | --- |
|
|
123
|
+
| <UTC timestamp> | <what was chosen> | <plain reason> | <SHA/PR/receipt/file:line> | <tests green/reverted/held/open> |
|
|
124
|
+
|
|
103
125
|
| Task | Dependencies | Owner | State | Revision evidence | Next action |
|
|
104
126
|
| --- | --- | --- | --- | --- | --- |
|
|
105
127
|
| <task> | <task IDs or none> | <role + session ID + worktree, or receipt ref> | <pending/in progress/complete/blocked> | <SHA + check/receipt refs> | <action + owner> |
|
|
@@ -41,7 +41,9 @@ This preserves the required contracts -> lifecycle -> audit load edge.
|
|
|
41
41
|
The current chat remains the driver under
|
|
42
42
|
[Standing contracts](../axstack/references/contracts.md). For each new
|
|
43
43
|
user round, the driver independently drafts the prioritized frontier and
|
|
44
|
-
recommendations
|
|
44
|
+
recommendations, except for an arena-grade question (below), where the driver
|
|
45
|
+
writes the brief and rubric but drafts no recommendation until the candidates
|
|
46
|
+
and judge verdicts return, so nothing anchors them. Then consult `axstack-advisor-astra` and
|
|
45
47
|
`axstack-advisor-fable` independently, without cross-reading, using the same
|
|
46
48
|
bounded evidence and question. Each adviser challenges assumptions, edges,
|
|
47
49
|
omissions, and alternatives; the driver synthesizes disagreements and accepts
|
|
@@ -56,6 +58,48 @@ unchanged receipts. Record compact adviser evidence, the driver's assessment,
|
|
|
56
58
|
and user-resolved choices for `axstack-spec`. If either adviser is unavailable,
|
|
57
59
|
hold Align; safe fact work may continue without substitution.
|
|
58
60
|
|
|
61
|
+
## Arena for hard-to-reverse design choices
|
|
62
|
+
|
|
63
|
+
Critique of one draft anchors every reader to that draft's shape. When a
|
|
64
|
+
question is arena-grade, the same test as for an ADR (a meaningful,
|
|
65
|
+
hard-to-reverse, non-obvious trade-off: architecture, module boundaries, data
|
|
66
|
+
model, migration strategy), replace the critique round for that question with
|
|
67
|
+
one arena round. Small or routine questions never enter the arena.
|
|
68
|
+
|
|
69
|
+
1. **Frame.** The driver writes the brief (the artifact, its constraints, the
|
|
70
|
+
settled decisions it must respect) and three to six gradeable rubric
|
|
71
|
+
criteria. Candidates receive only the brief; the rubric is for judging.
|
|
72
|
+
2. **Fan out.** `axstack-advisor-astra` and `axstack-advisor-fable` each
|
|
73
|
+
independently produce one candidate design plus a short rationale naming
|
|
74
|
+
the alternatives considered and rejected, from the same bounded evidence and
|
|
75
|
+
question, without cross-reading. The driver authors no candidate.
|
|
76
|
+
3. **Cross-judge.** After both candidates are complete, `axstack-arena-judge-astra`
|
|
77
|
+
and `axstack-arena-judge-fable` each independently score every candidate
|
|
78
|
+
per criterion from the rubric and candidates by label, and recommend a base
|
|
79
|
+
with a reason. Judges never author, never cross-read each other.
|
|
80
|
+
4. **Pick.** The driver reads every candidate end to end and scores per
|
|
81
|
+
criterion, not on holistic feel, then compares with both judges. Agreement
|
|
82
|
+
confirms the base. Disagreement between judges or with the driver means one
|
|
83
|
+
reading is biased or the rubric was ambiguous: re-read both rationales and
|
|
84
|
+
decide with a stated reason; never average verdicts or fabricate consensus.
|
|
85
|
+
5. **Graft.** Walk the losing candidate once more for the one or two ideas
|
|
86
|
+
worth porting and fold them into the base by hand so the result stays
|
|
87
|
+
coherent under one mental model. Convergence on the same shape is a strong
|
|
88
|
+
agreement signal: adopt the consensus shape, no graft. Wide divergence
|
|
89
|
+
means the frame was under-specified: reframe and rerun once, never
|
|
90
|
+
average.
|
|
91
|
+
6. **Present.** The synthesized design is the recommendation in the next
|
|
92
|
+
`Qn`, with its trade-off, judge verdicts, and what was grafted or rejected.
|
|
93
|
+
The user still decides; spec approval remains the one human checkpoint.
|
|
94
|
+
|
|
95
|
+
Record the synthesis note (base, grafts and their source candidate, rejections,
|
|
96
|
+
dropouts, both judge verdicts) as `Decisions` rows in the
|
|
97
|
+
[run record](../axstack/references/run-record.md). Load
|
|
98
|
+
[Orca runtime](../axstack/references/orca-runtime.md) immediately before the
|
|
99
|
+
first candidate or judge dispatch. If either adviser or judge seat is
|
|
100
|
+
unavailable, hold that question without substitution; unaffected fact work
|
|
101
|
+
and questions continue.
|
|
102
|
+
|
|
59
103
|
## Bound the interview
|
|
60
104
|
|
|
61
105
|
Twenty cumulative presented questions is the normal ceiling, not a target.
|
|
@@ -33,6 +33,18 @@ immediately before an actual profile dispatch.
|
|
|
33
33
|
4. For every gap, cite its inspected scope and applicable revision, stable
|
|
34
34
|
source identity, or content hash. “Not found” never means app-wide missing
|
|
35
35
|
without app-wide evidence; anything outside the inspected scope is unknown.
|
|
36
|
+
5. For "what could this break" or "blast radius of X", follow
|
|
37
|
+
[Blast radius](../axstack/references/blast-radius.md): find the breakage
|
|
38
|
+
beyond the diff and prove the one safety fact by running real code, or
|
|
39
|
+
mark it unproven.
|
|
40
|
+
6. For "show me your work" on existing work (a run, PR, branch, or change),
|
|
41
|
+
reconstruct the decision trail rather than re-describing the diff: read the
|
|
42
|
+
run record's `Decisions` and `Learnings`
|
|
43
|
+
([run record](../axstack/references/run-record.md)), then the commits, PR
|
|
44
|
+
body, review receipts, and ADRs. Present each consequential choice as what
|
|
45
|
+
was chosen, why, the evidence pointer, and its result; separate what the
|
|
46
|
+
record proves from what is inferred, and list choices with no recorded
|
|
47
|
+
reason as open rather than inventing one.
|
|
36
48
|
|
|
37
49
|
## 2. Choose proportional output
|
|
38
50
|
|
|
@@ -134,7 +134,9 @@ session is the owner for every PR it handles.
|
|
|
134
134
|
in authored mode the one reviewer covers all six angles:
|
|
135
135
|
1. Security and trust boundaries.
|
|
136
136
|
2. Correctness, failures, and edge cases.
|
|
137
|
-
3. Integration and regressions.
|
|
137
|
+
3. Integration and regressions: load [Blast radius](../axstack/references/blast-radius.md),
|
|
138
|
+
find what the change breaks beyond the diff, and grade the one fact it
|
|
139
|
+
is safe because of on the evidence ladder; below "ran it" is unproven.
|
|
138
140
|
4. Requirements, acceptance, and user behavior.
|
|
139
141
|
5. Architecture and solution design, including SOLID and credible simpler
|
|
140
142
|
alternatives.
|
|
@@ -233,6 +235,7 @@ Verdict: <APPROVE | REQUEST_CHANGES | INCOMPLETE>
|
|
|
233
235
|
Coverage: <angles + acceptance + executable evidence checked>
|
|
234
236
|
Limitations: <unverified boundaries + why>
|
|
235
237
|
Findings: <evidence + consequence each>
|
|
238
|
+
Safety fact: <the one fact the change is safe because of> — <ladder step + proof | unproven>
|
|
236
239
|
Escalate to user: <yes | no> — <criterion> — <reason>
|
|
237
240
|
```
|
|
238
241
|
|
package/src/roles.js
CHANGED
|
@@ -64,8 +64,8 @@ export function assessRoleReadiness(roles, preset) {
|
|
|
64
64
|
const gaps = [];
|
|
65
65
|
const isIntentionalAbsence = (role) => role.model === null && (
|
|
66
66
|
(preset === 'mixed' && role.id === 'axstack-checker') ||
|
|
67
|
-
(preset === 'codex-only' &&
|
|
68
|
-
(preset === 'claude-only' &&
|
|
67
|
+
(preset === 'codex-only' && ['axstack-advisor-fable', 'axstack-arena-judge-fable'].includes(role.id)) ||
|
|
68
|
+
(preset === 'claude-only' && ['axstack-advisor-astra', 'axstack-arena-judge-astra'].includes(role.id))
|
|
69
69
|
);
|
|
70
70
|
for (const role of roles) {
|
|
71
71
|
if (!bounds.has(role.provider)) {
|