axstack 0.21.0 → 0.22.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +3 -2
- package/docs/installation.md +10 -6
- package/docs/workflows.md +11 -9
- package/package.json +1 -1
- package/skills/axstack/references/automations.md +9 -0
- package/skills/axstack/references/contracts.md +3 -4
- package/skills/axstack/references/design-lens.md +3 -3
- package/skills/axstack/references/routing.md +7 -3
- package/skills/axstack/references/t3-runtime.md +3 -0
- package/skills/axstack/scripts/resolve-models.js +17 -9
- package/skills/axstack-align/SKILL.md +11 -58
- package/skills/axstack-brainstorm/SKILL.md +24 -0
- package/skills/axstack-brainstorm/references/arena.md +56 -0
package/README.md
CHANGED
|
@@ -14,7 +14,8 @@ coordination. You can start at the phase you need.
|
|
|
14
14
|
|
|
15
15
|
| Layer | Skill | What it does |
|
|
16
16
|
| --- | --- | --- |
|
|
17
|
-
| Plan | [axstack-align](skills/axstack-align/SKILL.md) | Settle scope through questions, a design lens sketch, and
|
|
17
|
+
| Plan | [axstack-align](skills/axstack-align/SKILL.md) | Settle scope through questions, a design lens sketch, and inline brainstorm validation. |
|
|
18
|
+
| Plan | [axstack-brainstorm](skills/axstack-brainstorm/SKILL.md) | Validate an approach with independent candidates; judges for hard choices. |
|
|
18
19
|
| Plan | [axstack-spec](skills/axstack-spec/SKILL.md) | Write and approve an observable specification. |
|
|
19
20
|
| Plan | [axstack-tickets](skills/axstack-tickets/SKILL.md) | Break approved scope into executable tasks. |
|
|
20
21
|
| Build | [axstack-implement](skills/axstack-implement/SKILL.md) | Build with strict TDD and an author → review → repair loop. |
|
|
@@ -36,7 +37,7 @@ implementation. Research, explanation, and peer review can start directly.
|
|
|
36
37
|
|
|
37
38
|
| Failure mode | How Axstack responds |
|
|
38
39
|
| --- | --- |
|
|
39
|
-
| Wrong thing built | Align rounds clarify the request; a
|
|
40
|
+
| Wrong thing built | Align rounds clarify the request; every brainstorm runs a light arena across configured families. Use judges only for Rung 2 hard-to-reverse choices. |
|
|
40
41
|
| Nobody really reviewed it | Strict TDD checks behavior first; with the mixed preset, cross-provider review checks the exact revision. |
|
|
41
42
|
| Design rot | The design lens sketches boundaries before a build; Improve surfaces evidenced changes later. |
|
|
42
43
|
| Agents left a mess | T3 makes delegation visible, one writer owns each PR, cleanup stays bounded, and a human merges. |
|
package/docs/installation.md
CHANGED
|
@@ -220,16 +220,20 @@ reconciles their findings.
|
|
|
220
220
|
|
|
221
221
|
The mixed checker and `axstack-research-web-google` have provider
|
|
222
222
|
`antigravity`; mixed `axstack-research-x` has provider `grok`. All three use
|
|
223
|
-
`model: null` with notes authorizing their agent-ID routes; T3 resolves
|
|
224
|
-
model from the first entry for that provider in saved capabilities.
|
|
225
|
-
Antigravity model
|
|
223
|
+
`model: null` with notes authorizing their agent-ID routes; T3 resolves Grok's
|
|
224
|
+
exact model from the first entry for that provider in saved capabilities.
|
|
225
|
+
For Antigravity `model:null` roles, select the first listed model whose ID ends
|
|
226
|
+
with `-<effort>` from saved capabilities and record its exact ID.
|
|
227
|
+
Missing effort-suffix matches hold resolution for Antigravity, including empty
|
|
228
|
+
model catalogs. The single-provider presets configure the checker and keep
|
|
226
229
|
both cross-provider research routes as intentional absences. Their
|
|
227
230
|
unavailable adviser and round-2 seat remain explicit same-provider
|
|
228
231
|
`model: null` roles, which do not make installation unready;
|
|
229
232
|
Align and Spec still hold until both Astra and Opus can return independent
|
|
230
|
-
receipts.
|
|
231
|
-
|
|
232
|
-
|
|
233
|
+
receipts. Every brainstorm runs a light arena across configured families.
|
|
234
|
+
Use judges only for Rung 2 hard-to-reverse choices: round 1 needs Opus; round 2,
|
|
235
|
+
if invoked, needs escalation Fable and Astra; a required seat that is unavailable
|
|
236
|
+
holds that round. The current chat drives on whatever
|
|
233
237
|
model runs it; no preset carries a driver role. Every other missing, invalid, unsupported, or unavailable role value holds only
|
|
234
238
|
the affected work. Codex and Claude class resolution reads the saved T3 capabilities catalog via
|
|
235
239
|
`skills/axstack/scripts/resolve-models.js --provider`; missing or malformed
|
package/docs/workflows.md
CHANGED
|
@@ -162,15 +162,17 @@ for explicit ownership transfer.
|
|
|
162
162
|
- `axstack-align` maps facts and dependencies, asks prioritized questions, and
|
|
163
163
|
consults Astra and Opus independently with the same bounded evidence and
|
|
164
164
|
question. It synthesizes disagreements and reuses unchanged receipts. For a
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
|
|
172
|
-
the
|
|
173
|
-
|
|
165
|
+
Rung 1 or 2 design question it loads `axstack-brainstorm` inline and reuses
|
|
166
|
+
its receipt instead of consulting twice; Align owns the interview.
|
|
167
|
+
- `axstack-brainstorm` validates an approach standalone or inline in the driver,
|
|
168
|
+
report-only. Every invocation compares independent Astra, Opus, Grok and
|
|
169
|
+
Antigravity candidates, including premise and smallest-change/do-nothing
|
|
170
|
+
checks. The driver scores, picks and grafts; Rung 1 uses no judges. At Rung 2,
|
|
171
|
+
`axstack-arena-judge-opus` judges round 1; disagreement on the base or caller
|
|
172
|
+
re-invocation with the user's rejection triggers fresh Fable/Astra round 2.
|
|
173
|
+
It returns a verdict, sketch and proposed questions, with no interview,
|
|
174
|
+
prototype or execution approval. Required seats hold; optional dropouts are
|
|
175
|
+
fenced. Fable also serves the bounded triggers in
|
|
174
176
|
[Standing contracts](../skills/axstack/references/contracts.md).
|
|
175
177
|
- `axstack-spec` writes observable acceptance, exclusions, decisions, and one
|
|
176
178
|
user-approved revision baseline. Linear through the executor MCP is the
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "axstack",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.22.0",
|
|
4
4
|
"description": "Axstack installer and setup CLI: installs owned chat skills and role data, configures supported harness settings, and checks T3 Code capabilities.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"claude-code",
|
|
@@ -271,6 +271,15 @@ The orphan sweep covers the run record's repositories plus registered repositori
|
|
|
271
271
|
`t3_thread_organize` settle/archive changes metadata only; exact guarded Git
|
|
272
272
|
worktree removal remains separate. Unknown, active or user-taken-over threads,
|
|
273
273
|
ambiguous publication and failed salvage stay preserved.
|
|
274
|
+
|
|
275
|
+
A review-manager pass may retire a predecessor pass's local `t3code/*` branch
|
|
276
|
+
with exact `git branch -D <branch>` only when ALL hold: the predecessor pass is
|
|
277
|
+
settled and proven run/lane-owned by recorded identity, its descendants and
|
|
278
|
+
evidence are settled, no worktree has the branch checked out, no remote
|
|
279
|
+
counterpart exists, its tip equals the recorded tip, and its tip is an ancestor
|
|
280
|
+
of the verified remote default branch. Otherwise hold; all other `git branch -d`
|
|
281
|
+
rules remain unchanged.
|
|
282
|
+
|
|
274
283
|
Past the authorized storage limit (default 20 retained lane worktrees), disable the schedule
|
|
275
284
|
with `update_scheduled_task` using `enabled:false` and hold.
|
|
276
285
|
Read back the disabled schedule with `list_scheduled_tasks`; uncertainty holds.
|
|
@@ -55,10 +55,9 @@ profile. Record the driver's provider and model in the run record.
|
|
|
55
55
|
|
|
56
56
|
For Align and Spec, the driver forms an independent assessment first, then
|
|
57
57
|
consults `axstack-advisor-astra` and `axstack-advisor-opus` independently with
|
|
58
|
-
the same bounded evidence and question. The
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
verdicts return. The driver synthesizes disagreements,
|
|
58
|
+
the same bounded evidence and question. The exception covers every brainstorm: the driver frames
|
|
59
|
+
its brief and rubric and assesses only after the candidates and required judge
|
|
60
|
+
rounds return. The driver synthesizes disagreements,
|
|
62
61
|
owns the decision, and the user still approves the spec. Reuse each valid
|
|
63
62
|
unchanged receipt; changed evidence, scope, or question requires a fresh
|
|
64
63
|
receipt. If either adviser is unavailable, Align and Spec hold without model or
|
|
@@ -15,15 +15,15 @@ sketch through the scope identity.
|
|
|
15
15
|
not reclassify work: Rung 1 can stay small. An unsettled material design
|
|
16
16
|
question still makes routing reassess size.
|
|
17
17
|
- **Rung 2 — arena.** A Rung 1 design that also meets the existing ADR test:
|
|
18
|
-
a meaningful, hard-to-reverse, non-obvious trade-off. Use
|
|
19
|
-
|
|
18
|
+
a meaningful, hard-to-reverse, non-obvious trade-off. Use [Brainstorm](../../axstack-brainstorm/SKILL.md)
|
|
19
|
+
with its Rung 2 judge rounds for that question.
|
|
20
20
|
|
|
21
21
|
There is no numeric threshold, file-count gate, or class-count gate.
|
|
22
22
|
|
|
23
23
|
## Questions in order
|
|
24
24
|
|
|
25
25
|
Ask only unresolved areas in numbered `Qn` rounds of one to three. Give each a
|
|
26
|
-
recommendation, reason, and trade-off; use Align's
|
|
26
|
+
recommendation, reason, and trade-off; use Align's inline brainstorm synthesis. Design and
|
|
27
27
|
arena questions share its unchanged budget: 20 normally, a justified extension
|
|
28
28
|
to 35, then opted-in refinement of at most five.
|
|
29
29
|
|
|
@@ -23,9 +23,11 @@ Resolve Codex and Claude classes with
|
|
|
23
23
|
`skills/axstack/scripts/resolve-models.js --provider <provider> --capabilities <path>`
|
|
24
24
|
using saved T3 capabilities JSON; missing or malformed capabilities holds.
|
|
25
25
|
Claude exact IDs come from capabilities, replacing transcript read-back.
|
|
26
|
-
Use an explicit model as given; for `model:null` without a class,
|
|
27
|
-
|
|
28
|
-
|
|
26
|
+
Use an explicit model as given; for `model:null` without a class, grok uses the
|
|
27
|
+
provider's first listed model from saved capabilities and records its exact ID.
|
|
28
|
+
For `model:null` without a class, Antigravity must select the first listed model
|
|
29
|
+
whose ID ends with `-<effort>` from saved capabilities and record its exact ID.
|
|
30
|
+
Missing effort-suffix matches hold resolution for Antigravity.
|
|
29
31
|
A Codex or Claude role with neither model nor class is an intentional absence
|
|
30
32
|
and holds.
|
|
31
33
|
Resume must reuse the saved capabilities and role snapshot with no re-resolution; changes require the user’s explicit decision.
|
|
@@ -50,6 +52,8 @@ step (3) for user routing: no substitution or same-provider review.
|
|
|
50
52
|
|
|
51
53
|
## Direct routes (no spec ceremony)
|
|
52
54
|
|
|
55
|
+
- Validate an approach -> `axstack-brainstorm`: inline, report-only independent
|
|
56
|
+
candidates; light arena always, judges only at Rung 2; return to the caller.
|
|
53
57
|
- Bounded research -> `axstack-research`: verify primary sources and code,
|
|
54
58
|
cite limits, and fan out distinct questions.
|
|
55
59
|
- Understand a system or gap -> `axstack-explain`:
|
|
@@ -32,6 +32,9 @@ A `model:null` role lacking a class must use the first model listed for its
|
|
|
32
32
|
provider in saved capabilities only for grok and antigravity (launch-by-agent-id
|
|
33
33
|
providers); record the exact ID, rather than an unresolved provider default.
|
|
34
34
|
Antigravity must hold when saved capabilities advertise zero models.
|
|
35
|
+
For Antigravity, first-listed selection must use the first model ID ending in `-<effort>` because its model ID encodes effort.
|
|
36
|
+
No matching effort suffix holds resolution for Antigravity.
|
|
37
|
+
For Antigravity, pass no effort option; effort read-back uses the model ID suffix.
|
|
35
38
|
For codex or claude, a role lacking both model and class is an intentional
|
|
36
39
|
absence and must hold; never use a provider default for that role.
|
|
37
40
|
|
|
@@ -45,7 +45,7 @@ function providerFrom(catalog, provider) {
|
|
|
45
45
|
return instance;
|
|
46
46
|
}
|
|
47
47
|
|
|
48
|
-
function modelFrom(models, { provider, model: pin, class: modelClass, excluded }) {
|
|
48
|
+
function modelFrom(models, { provider, model: pin, class: modelClass, excluded, effort }) {
|
|
49
49
|
if (pin && pin !== 'null') {
|
|
50
50
|
const model = models.find((model) => model.id === pin && !excluded.includes(model.id));
|
|
51
51
|
if (!model) throw new Error(`missing requested model ${pin}`);
|
|
@@ -56,8 +56,9 @@ function modelFrom(models, { provider, model: pin, class: modelClass, excluded }
|
|
|
56
56
|
if (['codex', 'claude'].includes(provider)) {
|
|
57
57
|
throw new Error(`intentional absence for ${provider}: missing model and class`);
|
|
58
58
|
}
|
|
59
|
-
//
|
|
60
|
-
const model =
|
|
59
|
+
// Antigravity encodes effort in its ID; exclusions never select a substitute.
|
|
60
|
+
const model = provider === 'antigravity'
|
|
61
|
+
? models.find((model) => model.id.endsWith(`-${effort}`)) : models[0];
|
|
61
62
|
if (!model || excluded.includes(model.id)) throw new Error('missing first listed model');
|
|
62
63
|
return model;
|
|
63
64
|
}
|
|
@@ -88,14 +89,21 @@ try {
|
|
|
88
89
|
const catalog = JSON.parse(readFileSync(path, 'utf8'));
|
|
89
90
|
const instance = providerFrom(catalog, provider);
|
|
90
91
|
const model = modelFrom(instance.models, options);
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
92
|
+
let effortOption = null;
|
|
93
|
+
if (provider === 'antigravity') {
|
|
94
|
+
if (!model.id.endsWith(`-${effort}`)) {
|
|
95
|
+
throw new Error(`unsupported effort ${effort} for ${provider}/${model.id}`);
|
|
96
|
+
}
|
|
97
|
+
} else {
|
|
98
|
+
const effortOptions = model.options.filter((option) => option.id === effortIds[provider]);
|
|
99
|
+
if (effortOptions.length !== 1 || !effortOptions[0].options?.some((value) => value.id === effort)
|
|
100
|
+
|| (provider === 'grok' && effort === 'max')) {
|
|
101
|
+
throw new Error(`unsupported effort ${effort} for ${provider}/${model.id}`);
|
|
102
|
+
}
|
|
103
|
+
effortOption = { id: effortOptions[0].id, value: effort };
|
|
96
104
|
}
|
|
97
105
|
console.log(JSON.stringify({ provider, providerInstanceId: instance.providerInstanceId,
|
|
98
|
-
model: model.id, effortOption
|
|
106
|
+
model: model.id, effortOption, source: 'capabilities', path }));
|
|
99
107
|
} catch (error) {
|
|
100
108
|
console.error(`model capabilities resolution hold: ${error.message}`);
|
|
101
109
|
process.exitCode = 1;
|
|
@@ -26,7 +26,9 @@ Set the rung from researched facts; never ask the user to choose it. A change
|
|
|
26
26
|
inside one module's existing interface, ownership, data flow, and failure
|
|
27
27
|
guarantees is Rung 0: no design questions or sketch. Otherwise load the
|
|
28
28
|
[design lens ladder](../axstack/references/design-lens.md) for Rung 1 or 2
|
|
29
|
-
and settle only unresolved areas in its order within the existing budget.
|
|
29
|
+
and settle only unresolved areas in its order within the existing budget.
|
|
30
|
+
For unresolved Rung 1 or 2 design questions, load
|
|
31
|
+
[Brainstorm](../axstack-brainstorm/SKILL.md) inline. Carry a
|
|
30
32
|
Rung 1 or 2 sketch in the substantial spec's `Design` section or the returned
|
|
31
33
|
small-change intent. A design question alone does not make small work
|
|
32
34
|
substantial; apply routing's existing size reassessment rule.
|
|
@@ -65,9 +67,10 @@ substantial; apply routing's existing size reassessment rule.
|
|
|
65
67
|
The current chat remains the driver under
|
|
66
68
|
[Standing contracts](../axstack/references/contracts.md). For each new
|
|
67
69
|
user round, the driver independently drafts the prioritized frontier and
|
|
68
|
-
recommendations, except for
|
|
69
|
-
|
|
70
|
-
|
|
70
|
+
recommendations, except for brainstorm questions, where the driver frames the
|
|
71
|
+
brief and rubric and assesses after the candidates and required judge rounds
|
|
72
|
+
return. Reuse valid brainstorm receipts to replace the adviser consult for
|
|
73
|
+
that question. For other questions, consult `axstack-advisor-astra` and
|
|
71
74
|
`axstack-advisor-opus` independently, without cross-reading, using the same
|
|
72
75
|
bounded evidence and question. Each adviser challenges assumptions, edges,
|
|
73
76
|
omissions, and alternatives; the driver synthesizes disagreements and accepts
|
|
@@ -87,60 +90,10 @@ the draft unchanged; it needs no new adviser pair. Changed draft text, a
|
|
|
87
90
|
blocking finding, or a high-stakes decision requires fresh receipts on the new
|
|
88
91
|
revision.
|
|
89
92
|
|
|
90
|
-
##
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
hard-to-reverse, non-obvious trade-off: architecture, module boundaries, data
|
|
95
|
-
model, migration strategy). Replace the critique round for that question with
|
|
96
|
-
an arena. Small or routine questions never enter the arena.
|
|
97
|
-
|
|
98
|
-
1. **Frame.** The driver writes the brief (the artifact, its constraints, the
|
|
99
|
-
settled decisions it must respect) and three to six gradeable rubric
|
|
100
|
-
criteria. Candidates receive only the brief; the rubric is for judging.
|
|
101
|
-
2. **Fan out.** Produce one candidate per configured family independently from the same brief,
|
|
102
|
-
without cross-reading: `axstack-advisor-astra`, `axstack-advisor-opus`,
|
|
103
|
-
`axstack-arena-candidate-grok`, and `axstack-arena-candidate-antigravity`.
|
|
104
|
-
Each gives a design, rationale, and rejected alternatives. The driver authors no candidate.
|
|
105
|
-
3. **Cross-judge.** After every candidate completes, give round 1 judge
|
|
106
|
-
`axstack-arena-judge-opus` the anonymized, relabeled candidates and rubric to
|
|
107
|
-
score every candidate per criterion and recommend a base with a reason.
|
|
108
|
-
The driver compares its own pick with the Opus verdict. Only if the driver
|
|
109
|
-
and the Opus judge disagree on the base, or the user rejects the round-1
|
|
110
|
-
synthesis,
|
|
111
|
-
run round 2 with fresh sessions: `axstack-escalation-fable` and `axstack-arena-judge-astra`
|
|
112
|
-
independently score the same anonymized candidates and rubric. Judges never
|
|
113
|
-
author, never cross-read each other, and never average verdicts. After round-2
|
|
114
|
-
verdicts return, the driver re-picks in step 4 and re-presents in step 6.
|
|
115
|
-
4. **Pick.** The driver reads every candidate end to end and scores per
|
|
116
|
-
criterion, not on holistic feel, then compares with the judge verdicts from
|
|
117
|
-
each completed round. Agreement confirms the base. On disagreement, re-read
|
|
118
|
-
the rationales and decide with a stated reason; never average verdicts or
|
|
119
|
-
fabricate consensus.
|
|
120
|
-
5. **Graft.** Walk the losing candidates once more for the one or two ideas
|
|
121
|
-
worth porting and fold them into the base by hand so the result stays
|
|
122
|
-
coherent under one mental model. Convergence on the same shape is a strong
|
|
123
|
-
agreement signal: adopt the consensus shape, no graft. Wide divergence
|
|
124
|
-
means the frame was under-specified: reframe and rerun once, never
|
|
125
|
-
average.
|
|
126
|
-
6. **Present.** The synthesized design is the recommendation in the next
|
|
127
|
-
`Qn`, with its trade-off, judge verdicts per round, and what was grafted or rejected.
|
|
128
|
-
The user still decides; spec approval remains the one human checkpoint.
|
|
129
|
-
|
|
130
|
-
Record the synthesis note (base, grafts and their source candidate, rejections,
|
|
131
|
-
dropouts, judge verdicts per round) as `Decisions` rows in the
|
|
132
|
-
[run record](../axstack/references/run-record.md). Load
|
|
133
|
-
[T3 runtime](../axstack/references/t3-runtime.md) immediately before the
|
|
134
|
-
first candidate or judge dispatch. If an optional Grok or Antigravity candidate
|
|
135
|
-
malfunctions (launch failure, trust/login prompt, or prompt block), fence it,
|
|
136
|
-
record `absent (<reason>)`, name it once in the next
|
|
137
|
-
read-back, and continue with available candidates without relay or substitution.
|
|
138
|
-
A required adviser, candidate, or judge unavailable at launch or returning a
|
|
139
|
-
failed receipt holds that question without substitution; record the gap and ask
|
|
140
|
-
whether to proceed. In mixed fan-out retain at least one Codex and one Claude
|
|
141
|
-
seat, or hold the affected question.
|
|
142
|
-
For an uncertain dispatch, reconcile natively; it is never treated as absent.
|
|
143
|
-
Unaffected fact work and questions continue.
|
|
93
|
+
## Use the brainstorm synthesis
|
|
94
|
+
|
|
95
|
+
Present its recommendation and trade-off in the next `Qn` within the same
|
|
96
|
+
question budget. The user still decides; reuse unchanged receipts.
|
|
144
97
|
|
|
145
98
|
## Bound the interview
|
|
146
99
|
|
|
@@ -0,0 +1,24 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: axstack-brainstorm
|
|
3
|
+
description: When validating a design approach, use axstack-brainstorm to compare independent candidates and return a synthesis.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Brainstorm
|
|
7
|
+
|
|
8
|
+
Standalone usage: `/axstack-brainstorm <problem + candidate approach>`.
|
|
9
|
+
Run inline in the current T3 driver thread. Never delegate a coordinator.
|
|
10
|
+
This procedure is report-only. Do not implement, prototype, or grant execution
|
|
11
|
+
or spec approval. Never interview the user or invoke Align.
|
|
12
|
+
|
|
13
|
+
Load [Standing contracts](../axstack/references/contracts.md) before acting.
|
|
14
|
+
Set the rung from researched facts under the
|
|
15
|
+
[design lens](../axstack/references/design-lens.md); Rung 2 meets its ADR test.
|
|
16
|
+
Follow the [arena procedure](references/arena.md), preserving settled decisions
|
|
17
|
+
and naming missing evidence. Candidates challenge the supplied approach as
|
|
18
|
+
well as proposing alternatives.
|
|
19
|
+
|
|
20
|
+
Return `recommend | revise | reject | unresolved` with the
|
|
21
|
+
[sketch](../axstack/references/design-lens.md#sketch), including `Rejected` and
|
|
22
|
+
`Open`, plus base, criterion scores, grafts, dropouts and completed judge rounds.
|
|
23
|
+
Return proposed questions to the caller. The caller owns preferences, next
|
|
24
|
+
steps and any approval; an open blocker stays unresolved.
|
|
@@ -0,0 +1,56 @@
|
|
|
1
|
+
## Arena
|
|
2
|
+
|
|
3
|
+
Every brainstorm runs a light arena. At Rung 1, the driver scores, picks and
|
|
4
|
+
grafts without a judge round. Only at Rung 2, run the judge rounds below for
|
|
5
|
+
hard-to-reverse choices; they replace the critique round for that question.
|
|
6
|
+
|
|
7
|
+
1. **Frame.** The driver writes the brief (the artifact, its constraints, the
|
|
8
|
+
settled decisions it must respect) and three to six gradeable rubric
|
|
9
|
+
criteria. Every brief requires a premise check and comparison with the
|
|
10
|
+
smallest change and doing nothing. The rubric uses agent-contributor red
|
|
11
|
+
flags as evidence prompts: seeing only opened files, copying the nearest
|
|
12
|
+
example, taking the shortest path that compiles. Treat them as prompts, not defects. Candidates receive only the brief; the rubric is for judging.
|
|
13
|
+
2. **Fan out.** Produce one candidate per configured family independently from the same brief,
|
|
14
|
+
without cross-reading: `axstack-advisor-astra`, `axstack-advisor-opus`,
|
|
15
|
+
`axstack-arena-candidate-grok`, and `axstack-arena-candidate-antigravity`.
|
|
16
|
+
Each gives a design, rationale, and rejected alternatives. The driver authors no candidate.
|
|
17
|
+
The driver drafts no recommendation until every candidate returns.
|
|
18
|
+
3. **Cross-judge (Rung 2 only).** After every candidate completes, give round 1 judge
|
|
19
|
+
`axstack-arena-judge-opus` the anonymized, relabeled candidates and rubric to
|
|
20
|
+
score every candidate per criterion and recommend a base with a reason.
|
|
21
|
+
The driver compares its own pick with the Opus verdict. Only if the driver
|
|
22
|
+
and the Opus judge disagree on the base, or the caller re-invokes with the user's rejection of the round-1
|
|
23
|
+
synthesis,
|
|
24
|
+
run round 2 with fresh sessions: `axstack-escalation-fable` and `axstack-arena-judge-astra`
|
|
25
|
+
independently score the same anonymized candidates and rubric. Judges never
|
|
26
|
+
author, never cross-read each other, and never average verdicts. After round-2
|
|
27
|
+
verdicts return, the driver re-picks in step 4 and re-presents in step 6.
|
|
28
|
+
4. **Pick.** The driver reads every candidate end to end and scores per
|
|
29
|
+
criterion, not on holistic feel, then compares with the judge verdicts from
|
|
30
|
+
each completed round. Agreement confirms the base. On disagreement, re-read
|
|
31
|
+
the rationales and decide with a stated reason; never average verdicts or
|
|
32
|
+
fabricate consensus.
|
|
33
|
+
5. **Graft.** Revisit losing candidates once; graft one or two ideas by hand
|
|
34
|
+
into the coherent base under one mental model. Convergence on the same shape is a strong
|
|
35
|
+
agreement signal: adopt the consensus shape, no graft. Wide divergence
|
|
36
|
+
means the frame was under-specified: reframe and rerun once, never
|
|
37
|
+
average.
|
|
38
|
+
6. **Present.** Return the synthesized design to the caller with its
|
|
39
|
+
trade-off, judge verdicts per round, and what was grafted or rejected.
|
|
40
|
+
The user still decides; spec approval remains the one human checkpoint.
|
|
41
|
+
|
|
42
|
+
Record the synthesis note (base, graft sources, rejections, dropouts, judge
|
|
43
|
+
verdicts per round) as `Decisions` rows in the
|
|
44
|
+
[caller's run record](../../axstack/references/run-record.md) when one exists,
|
|
45
|
+
otherwise include it in the returned verdict. Load
|
|
46
|
+
[T3 runtime](../../axstack/references/t3-runtime.md) immediately before the
|
|
47
|
+
first candidate or judge dispatch. If an optional Grok or Antigravity candidate
|
|
48
|
+
malfunctions (launch failure, trust/login prompt, or prompt block), fence it,
|
|
49
|
+
record `absent (<reason>)`, name it once in the next
|
|
50
|
+
read-back, and continue with available candidates without relay or substitution.
|
|
51
|
+
A required adviser, candidate, or judge needed for this rung unavailable at
|
|
52
|
+
launch or returning a failed receipt holds that question without substitution;
|
|
53
|
+
record the gap and return a proposed question to the caller about whether to proceed. In mixed fan-out retain at least one Codex and one Claude
|
|
54
|
+
seat, or hold the affected question.
|
|
55
|
+
For an uncertain dispatch, reconcile natively; it is never treated as absent.
|
|
56
|
+
Unaffected fact work and questions continue.
|