axstack 0.21.0 → 0.23.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -14,7 +14,8 @@ coordination. You can start at the phase you need.
14
14
 
15
15
  | Layer | Skill | What it does |
16
16
  | --- | --- | --- |
17
- | Plan | [axstack-align](skills/axstack-align/SKILL.md) | Settle scope through questions, a design lens sketch, and a four-family arena for hard choices. |
17
+ | Plan | [axstack-align](skills/axstack-align/SKILL.md) | Settle scope through questions, a design lens sketch, and inline brainstorm validation. |
18
+ | Plan | [axstack-brainstorm](skills/axstack-brainstorm/SKILL.md) | Validate an approach with independent candidates; judges for hard choices. |
18
19
  | Plan | [axstack-spec](skills/axstack-spec/SKILL.md) | Write and approve an observable specification. |
19
20
  | Plan | [axstack-tickets](skills/axstack-tickets/SKILL.md) | Break approved scope into executable tasks. |
20
21
  | Build | [axstack-implement](skills/axstack-implement/SKILL.md) | Build with strict TDD and an author → review → repair loop. |
@@ -22,6 +23,7 @@ coordination. You can start at the phase you need.
22
23
  | Verify | [axstack-review](skills/axstack-review/SKILL.md) | Review a PR or bounded codebase at an exact revision. |
23
24
  | Verify | [axstack-improve](skills/axstack-improve/SKILL.md) | Find evidenced codebase improvements without editing code. |
24
25
  | Verify | [axstack-audit](skills/axstack-audit/SKILL.md) | Measure a run's outcomes and evidence gaps. |
26
+ | Verify | [axstack-correct](skills/axstack-correct/SKILL.md) | Report repeated mistakes and propose stronger checks when invoked by the user. |
25
27
  | Operate | [axstack-watch](skills/axstack-watch/SKILL.md) | Observe or maintain an existing PR within its authority. |
26
28
  | Operate | [axstack-cleanup](skills/axstack-cleanup/SKILL.md) | Retire eligible completed agent resources. |
27
29
  | Operate | [axstack-relay](skills/axstack-relay/SKILL.md) | Send an explicit message or authorized notification. |
@@ -36,7 +38,7 @@ implementation. Research, explanation, and peer review can start directly.
36
38
 
37
39
  | Failure mode | How Axstack responds |
38
40
  | --- | --- |
39
- | Wrong thing built | Align rounds clarify the request; a four-family arena compares approaches for hard choices. |
41
+ | Wrong thing built | Align rounds clarify the request; every brainstorm runs a light arena across configured families. Use judges only for Rung 2 hard-to-reverse choices. |
40
42
  | Nobody really reviewed it | Strict TDD checks behavior first; with the mixed preset, cross-provider review checks the exact revision. |
41
43
  | Design rot | The design lens sketches boundaries before a build; Improve surfaces evidenced changes later. |
42
44
  | Agents left a mess | T3 makes delegation visible, one writer owns each PR, cleanup stays bounded, and a human merges. |
@@ -220,16 +220,20 @@ reconciles their findings.
220
220
 
221
221
  The mixed checker and `axstack-research-web-google` have provider
222
222
  `antigravity`; mixed `axstack-research-x` has provider `grok`. All three use
223
- `model: null` with notes authorizing their agent-ID routes; T3 resolves the exact
224
- model from the first entry for that provider in saved capabilities. Empty
225
- Antigravity model catalogs hold. The single-provider presets configure the checker and keep
223
+ `model: null` with notes authorizing their agent-ID routes; T3 resolves Grok's
224
+ exact model from the first entry for that provider in saved capabilities.
225
+ For Antigravity `model:null` roles, select the first listed model whose ID ends
226
+ with `-<effort>` from saved capabilities and record its exact ID.
227
+ Missing effort-suffix matches hold resolution for Antigravity, including empty
228
+ model catalogs. The single-provider presets configure the checker and keep
226
229
  both cross-provider research routes as intentional absences. Their
227
230
  unavailable adviser and round-2 seat remain explicit same-provider
228
231
  `model: null` roles, which do not make installation unready;
229
232
  Align and Spec still hold until both Astra and Opus can return independent
230
- receipts. For an arena-grade Align question, round 1 needs Opus; round 2, if
231
- invoked, needs escalation Fable and Astra; a required seat that is unavailable holds that
232
- round. The current chat drives on whatever
233
+ receipts. Every brainstorm runs a light arena across configured families.
234
+ Use judges only for Rung 2 hard-to-reverse choices: round 1 needs Opus; round 2,
235
+ if invoked, needs escalation Fable and Astra; a required seat that is unavailable
236
+ holds that round. The current chat drives on whatever
233
237
  model runs it; no preset carries a driver role. Every other missing, invalid, unsupported, or unavailable role value holds only
234
238
  the affected work. Codex and Claude class resolution reads the saved T3 capabilities catalog via
235
239
  `skills/axstack/scripts/resolve-models.js --provider`; missing or malformed
package/docs/workflows.md CHANGED
@@ -18,6 +18,8 @@ and PR shape.
18
18
  Direct routes need no spec ceremony:
19
19
 
20
20
  - `axstack-research` answers one bounded source-backed question.
21
+ - `axstack-correct` reports repeated mistakes and proposes stronger checks.
22
+ Only the user invokes it.
21
23
  - `axstack-explain` separates implemented, intended, tested, live, and unknown
22
24
  behavior; complex visuals receive exact-artifact QA where applicable.
23
25
  - `axstack-improve` returns a small ranked set of evidenced improvement
@@ -162,15 +164,17 @@ for explicit ownership transfer.
162
164
  - `axstack-align` maps facts and dependencies, asks prioritized questions, and
163
165
  consults Astra and Opus independently with the same bounded evidence and
164
166
  question. It synthesizes disagreements and reuses unchanged receipts. For a
165
- hard-to-reverse design choice it runs an arena instead: Astra, Opus, Grok,
166
- and Antigravity each author a candidate. `axstack-arena-judge-opus` scores
167
- them in round 1; the driver compares its own pick with that verdict. If they
168
- disagree on the base or the user rejects the round-1 synthesis,
169
- `axstack-escalation-fable` and `axstack-arena-judge-astra` independently
170
- score the same anonymized candidates and rubric in round 2. The driver
171
- picks a base, grafts strong ideas, and records judge verdicts per round in
172
- the `Decisions` rows without averaging. Fable escalation uses a fresh
173
- session for round 2, high-stakes agreement, or the bounded trigger in
167
+ Rung 1 or 2 design question it loads `axstack-brainstorm` inline and reuses
168
+ its receipt instead of consulting twice; Align owns the interview.
169
+ - `axstack-brainstorm` validates an approach standalone or inline in the driver,
170
+ report-only. Every invocation compares independent Astra, Opus, Grok and
171
+ Antigravity candidates, including premise and smallest-change/do-nothing
172
+ checks. The driver scores, picks and grafts; Rung 1 uses no judges. At Rung 2,
173
+ `axstack-arena-judge-opus` judges round 1; disagreement on the base or caller
174
+ re-invocation with the user's rejection triggers fresh Fable/Astra round 2.
175
+ It returns a verdict, sketch and proposed questions, with no interview,
176
+ prototype or execution approval. Required seats hold; optional dropouts are
177
+ fenced. Fable also serves the bounded triggers in
174
178
  [Standing contracts](../skills/axstack/references/contracts.md).
175
179
  - `axstack-spec` writes observable acceptance, exclusions, decisions, and one
176
180
  user-approved revision baseline. Linear through the executor MCP is the
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "axstack",
3
- "version": "0.21.0",
3
+ "version": "0.23.0",
4
4
  "description": "Axstack installer and setup CLI: installs owned chat skills and role data, configures supported harness settings, and checks T3 Code capabilities.",
5
5
  "keywords": [
6
6
  "claude-code",
@@ -271,6 +271,15 @@ The orphan sweep covers the run record's repositories plus registered repositori
271
271
  `t3_thread_organize` settle/archive changes metadata only; exact guarded Git
272
272
  worktree removal remains separate. Unknown, active or user-taken-over threads,
273
273
  ambiguous publication and failed salvage stay preserved.
274
+
275
+ A review-manager pass may retire a predecessor pass's local `t3code/*` branch
276
+ with exact `git branch -D <branch>` only when ALL hold: the predecessor pass is
277
+ settled and proven run/lane-owned by recorded identity, its descendants and
278
+ evidence are settled, no worktree has the branch checked out, no remote
279
+ counterpart exists, its tip equals the recorded tip, and its tip is an ancestor
280
+ of the verified remote default branch. Otherwise hold; all other `git branch -d`
281
+ rules remain unchanged.
282
+
274
283
  Past the authorized storage limit (default 20 retained lane worktrees), disable the schedule
275
284
  with `update_scheduled_task` using `enabled:false` and hold.
276
285
  Read back the disabled schedule with `list_scheduled_tasks`; uncertainty holds.
@@ -17,7 +17,16 @@ protection; a mismatch holds publication. If the push outcome is ambiguous,
17
17
  inspect remote state before retrying.
18
18
 
19
19
  Before reviewer dispatch, read the remote ref back and confirm that it resolves
20
- to the candidate SHA; also pin the current base. Record:
20
+ to the candidate SHA; also pin the current base.
21
+ After every push, before post-push diligence or merge-ready, compare the PR body's
22
+ stated head SHA with `Confirmed remote SHA` and record `PR body head SHA: <sha or none>`.
23
+ Accept `none` for a PR body without a stated head SHA.
24
+ Accept a stated full head SHA only when it equals `Confirmed remote SHA`.
25
+ Accept a stated abbreviated head SHA only when it is a matching prefix of
26
+ `Confirmed remote SHA` with at least 7 hexadecimal characters.
27
+ Hold on a mismatched head SHA or an abbreviation shorter than 7 characters.
28
+ Leave other SHAs in the PR body unchanged.
29
+ Record:
21
30
 
22
31
  ```text
23
32
  Candidate: <sha>
@@ -25,6 +34,7 @@ Base: <sha>
25
34
  Remote ref: <branch>
26
35
  Expected-old remote SHA: <sha | absent>
27
36
  Confirmed remote SHA: <sha>
37
+ PR body head SHA: <sha or none>
28
38
  PR: <url>
29
39
  CI: <run ID or URL and triggered/pending/completed status>
30
40
  ```
@@ -2,10 +2,12 @@
2
2
 
3
3
  Apply these authority, scope, and model rules before consequential action.
4
4
 
5
+ For user-facing output, follow [STE-inspired writing](ste-writing.md).
6
+
5
7
  ## Required lifecycle load
6
8
 
7
- Except for `axstack-audit` and `axstack-relay`, every independently called phase
8
- must load and follow [Shared lifecycle](lifecycle.md) before acting. When a
9
+ Except for `axstack-audit`, `axstack-relay`, and `axstack-correct`, every
10
+ independently called phase must load and follow [Shared lifecycle](lifecycle.md) before acting. When a
9
11
  substantive run ends or reaches a meaningful checkpoint, apply the lifecycle
10
12
  audit hook. The audit phase loads these contracts, writes its assigned record,
11
13
  and stops; it never audits itself.
@@ -55,10 +57,9 @@ profile. Record the driver's provider and model in the run record.
55
57
 
56
58
  For Align and Spec, the driver forms an independent assessment first, then
57
59
  consults `axstack-advisor-astra` and `axstack-advisor-opus` independently with
58
- the same bounded evidence and question. The one exception is an arena-grade
59
- Align question: there the driver frames the brief and rubric, the advisers
60
- author candidates, and the driver assesses only after the candidates and judge
61
- verdicts return. The driver synthesizes disagreements,
60
+ the same bounded evidence and question. The exception covers every brainstorm: the driver frames
61
+ its brief and rubric and assesses only after the candidates and required judge
62
+ rounds return. The driver synthesizes disagreements,
62
63
  owns the decision, and the user still approves the spec. Reuse each valid
63
64
  unchanged receipt; changed evidence, scope, or question requires a fresh
64
65
  receipt. If either adviser is unavailable, Align and Spec hold without model or
@@ -15,15 +15,15 @@ sketch through the scope identity.
15
15
  not reclassify work: Rung 1 can stay small. An unsettled material design
16
16
  question still makes routing reassess size.
17
17
  - **Rung 2 — arena.** A Rung 1 design that also meets the existing ADR test:
18
- a meaningful, hard-to-reverse, non-obvious trade-off. Use Align's all-family
19
- arena for that question.
18
+ a meaningful, hard-to-reverse, non-obvious trade-off. Use [Brainstorm](../../axstack-brainstorm/SKILL.md)
19
+ with its Rung 2 judge rounds for that question.
20
20
 
21
21
  There is no numeric threshold, file-count gate, or class-count gate.
22
22
 
23
23
  ## Questions in order
24
24
 
25
25
  Ask only unresolved areas in numbered `Qn` rounds of one to three. Give each a
26
- recommendation, reason, and trade-off; use Align's adviser critique. Design and
26
+ recommendation, reason, and trade-off; use Align's inline brainstorm synthesis. Design and
27
27
  arena questions share its unchanged budget: 20 normally, a justified extension
28
28
  to 35, then opted-in refinement of at most five.
29
29
 
@@ -3,7 +3,7 @@
3
3
  Read the [T3 runtime boundary](t3-runtime.md) before dispatch.
4
4
  Dispatch `axstack-diligence` with async `delegate_task` in the driver worktree,
5
5
  using a pinned brief and private evidence paths.
6
- It is read-only, never authors or edits, and returns `PASS` or `FINDINGS`
6
+ It is read-only, never authors or edits, and returns `PASS`, `FINDINGS`, or `UNKNOWN`
7
7
  with locations, observed evidence, and limits. A stale or missing receipt is
8
8
  not a pass. Keep its first pass independent of other reviewers and workers.
9
9
 
@@ -21,5 +21,14 @@ publication, compare the author receipt with its evidence folder: red/green
21
21
  logs exist, and counts, SHAs, and paths match. For release preparation, compare
22
22
  the release PR body with the merged PRs.
23
23
 
24
+ Retain the full output of every full-suite run as a named log in the dispatch's
25
+ private evidence folder.
26
+ For an unattributed full-suite failure, report `UNKNOWN` and return the failure to the driver.
27
+ Never report `PASS` for an unattributed full-suite failure.
28
+ The driver records the disposition of each full-suite failure in the run record.
29
+ Report an attributed failure with its retained output normally.
30
+ An observed full-suite failure still fails the suite, including when attributed or reported `UNKNOWN`.
31
+ No reruns are required.
32
+
24
33
  `FINDINGS` identifies a mismatch for the driver to resolve at the owning phase;
25
34
  it does not edit the artifact or create another review round by itself.
@@ -118,6 +118,6 @@ After preservation and salvage checks, archive the exact eligible T3 thread
118
118
  with `t3_thread_organize`, then remove its exact recorded checkout path using
119
119
  `git worktree remove <path>` without force. Verify absence with
120
120
  `git worktree list --porcelain`; retire only eligible local-only branches with
121
- `git branch -d` under workspace hygiene. Never use shell recursive deletion or
121
+ `git branch -d` under workspace hygiene. Never use shell recursive deletion to remove a whole worktree or
122
122
  treat archive success as ownership, settlement, liveness or cleanup proof.
123
123
  Failure or uncertainty preserves the resource.
@@ -143,6 +143,8 @@ nothing without tested independent review.
143
143
 
144
144
  ## Close-out
145
145
 
146
+ Close-out requires the [close-out acceptance table](run-record.md#close-out-acceptance).
147
+
146
148
  PRs merge by forge state; close out: (1) settle every T3 worker run and archive eligible threads; (2) compact record with counts and denominators—user
147
149
  interventions/deviations from plan/repairs; (3) `axstack-auditor`: settle
148
150
  non-zero/requested, else `counts zero`. A base auditor preflight rejection
@@ -153,7 +155,7 @@ use `axstack-cleanup`, remove the run's own scheduled tasks under
153
155
  [Workspace hygiene](workspace-hygiene.md), and close selected external-tracker tickets;
154
156
  (5) mark the
155
157
  [Run record](run-record.md) `Archived`. `Archived`—one each:
156
- settlement receipt; compact record path; auditor decision plus settlement
158
+ settlement receipt; compact record path; close-out acceptance table; auditor decision plus settlement
157
159
  receipt, `counts zero`, or the unlaunchable UNKNOWN archive receipt; scheduled-task,
158
160
  release, and ticket receipts; archive timestamp.
159
161
  `active`/receipt-incomplete record: close-out pending, never done. One-step
@@ -0,0 +1,22 @@
1
+ # Performance checklist
2
+
3
+ Use this checklist only for performance claims.
4
+ Keep workflow count accounting with axstack-audit.
5
+ A limiter is the resource or code path that bounds performance.
6
+
7
+ 1. **Limiter: What bounds the result?**
8
+ Identify the resource or code path that limits the result from profiles or counters captured during a run.
9
+ 2. **Tuning: Did each option use suitable settings?**
10
+ Check each option with comparable production settings, versions, and data.
11
+ 3. **Limits: Is the result physically plausible?**
12
+ Compare the result with hardware limits and the changed code's share of total elapsed time.
13
+ 4. **Errors: Did the run succeed?**
14
+ Verify successful work and correct outputs using error counts from the measurement run.
15
+ 5. **Reproducibility: Does the difference repeat?**
16
+ Report the median and range from repeated, alternating runs of each option.
17
+ 6. **Relevance: Does the result matter to users?**
18
+ Measure the end-to-end user path with realistic data and concurrency.
19
+ 7. **Work happened: Did the timed work finish?**
20
+ Verify that the intended work completed inside the timed region.
21
+
22
+ Ideas paraphrased from [pstack benchmark-checklist](https://github.com/cursor/plugins/blob/e43c7ee26e0038c6c1fa8380dd34ce86ff94cb2a/pstack/skills/benchmark-checklist/SKILL.md) (MIT), using Brendan Gregg's seven benchmark questions.
@@ -23,9 +23,11 @@ Resolve Codex and Claude classes with
23
23
  `skills/axstack/scripts/resolve-models.js --provider <provider> --capabilities <path>`
24
24
  using saved T3 capabilities JSON; missing or malformed capabilities holds.
25
25
  Claude exact IDs come from capabilities, replacing transcript read-back.
26
- Use an explicit model as given; for `model:null` without a class, only grok and
27
- antigravity use the provider's first listed model from saved capabilities and
28
- record its exact ID.
26
+ Use an explicit model as given; for `model:null` without a class, grok uses the
27
+ provider's first listed model from saved capabilities and records its exact ID.
28
+ For `model:null` without a class, Antigravity must select the first listed model
29
+ whose ID ends with `-<effort>` from saved capabilities and record its exact ID.
30
+ Missing effort-suffix matches hold resolution for Antigravity.
29
31
  A Codex or Claude role with neither model nor class is an intentional absence
30
32
  and holds.
31
33
  Resume must reuse the saved capabilities and role snapshot with no re-resolution; changes require the user’s explicit decision.
@@ -50,6 +52,8 @@ step (3) for user routing: no substitution or same-provider review.
50
52
 
51
53
  ## Direct routes (no spec ceremony)
52
54
 
55
+ - Validate an approach -> `axstack-brainstorm`: inline, report-only independent
56
+ candidates; light arena always, judges only at Rung 2; return to the caller.
53
57
  - Bounded research -> `axstack-research`: verify primary sources and code,
54
58
  cite limits, and fan out distinct questions.
55
59
  - Understand a system or gap -> `axstack-explain`:
@@ -61,6 +65,8 @@ step (3) for user routing: no substitution or same-provider review.
61
65
  off a classified repair (explain: how; debug: what's wrong).
62
66
  - Code quality/refactor discovery -> `axstack-improve`: rank bounded
63
67
  candidates with evidence; report only, no source edits.
68
+ - Repeated mistakes need evidence and stronger checks -> `axstack-correct`:
69
+ user-invoked, report only.
64
70
  - Accepted worker/task/run completion or bounded backlog request -> driver invokes
65
71
  `axstack-cleanup` inline; never dispatch it.
66
72
  - Preparation completion, watch expiry, resume, or reconciliation -> the
@@ -119,6 +119,21 @@ by the next owner:
119
119
  evidence lives. Read it on resume before reconciling; append, never rewrite,
120
120
  and keep entries as short as the evidence pointer allows.
121
121
 
122
+ ## Close-out acceptance
123
+
124
+ Use one row per acceptance check, including each clause of a compound check.
125
+ Record in each row a passing evidence pointer or a user-accepted hold with its `Decisions` row.
126
+ If neither passing evidence nor a `Decisions` row with a user-accepted hold exists for an acceptance clause, hold close-out.
127
+ A recorded user-accepted hold in `Decisions` satisfies that clause for close-out; keep the unmet result explicit.
128
+ Record the reason for each driver-elected repair in `Decisions`.
129
+
130
+ ```markdown
131
+ | Acceptance check | Result | Evidence or user-accepted hold | Repair election (Decisions row + reason, or none) |
132
+ | --- | --- | --- | --- |
133
+ | <check + clause> | passed | <SHA + check/log pointer> | <Decisions row + reason, or none> |
134
+ | <check + unmet clause> | held (user accepted) | <Decisions row + user acceptance receipt> | <Decisions row + reason, or none> |
135
+ ```
136
+
122
137
  ## Privacy
123
138
 
124
139
  Record concise IDs, SHAs, URLs, status, timestamps, next actions, and evidence
@@ -0,0 +1,22 @@
1
+ # STE-inspired writing
2
+
3
+ STE means Simplified Technical English.
4
+ This reference borrows clarity rules from STE.
5
+
6
+ Follow these rules only for new or materially revised user-facing output.
7
+ User-facing output includes reports, PR bodies, briefs, read-backs, and packaged guidance.
8
+ Do not rewrite existing prose for style.
9
+ Keep exact identifiers unchanged.
10
+ Keep quotes unchanged.
11
+ Keep safety contracts unchanged.
12
+ Keep prompt-byte contracts unchanged.
13
+
14
+ Give each sentence one instruction.
15
+ Keep sentences short.
16
+ Write in active voice.
17
+ If a step has a condition, state the condition before the step.
18
+ Define technical terms before their first use.
19
+ Use one name consistently for each term.
20
+ Use words instead of slashes for "and" or "or".
21
+
22
+ Ideas drawn from [pstack technical-writing](https://github.com/cursor/plugins/blob/e43c7ee26e0038c6c1fa8380dd34ce86ff94cb2a/pstack/skills/technical-writing/SKILL.md#L65-L88) (MIT).
@@ -32,6 +32,9 @@ A `model:null` role lacking a class must use the first model listed for its
32
32
  provider in saved capabilities only for grok and antigravity (launch-by-agent-id
33
33
  providers); record the exact ID, rather than an unresolved provider default.
34
34
  Antigravity must hold when saved capabilities advertise zero models.
35
+ For Antigravity, first-listed selection must use the first model ID ending in `-<effort>` because its model ID encodes effort.
36
+ No matching effort suffix holds resolution for Antigravity.
37
+ For Antigravity, pass no effort option; effort read-back uses the model ID suffix.
35
38
  For codex or claude, a role lacking both model and class is an intentional
36
39
  absence and must hold; never use a provider default for that role.
37
40
 
@@ -14,7 +14,7 @@ Validate its real path, absence of symlinks and ownership before use and cleanup
14
14
  Evidence files still go to the private `<run>/evidence/<key>/` folder.
15
15
 
16
16
  Every shell deletion targets a literal absolute path or a `${VAR:?}`-guarded expansion, only inside the worker's own evidence folder, `TMPDIR`, or worktree.
17
- For example, `rm -rf -- "${EV:?}/mut"` requires a validated owned evidence path.
17
+ For validated owned scratch, use `rm -r /tmp/<dispatch-key>/scratch` on a literal absolute path inside the evidence folder, `TMPDIR`, or worktree.
18
18
  Never use a bare `$VAR`, a glob on a variable, `/`, `HOME`, or a shared root as a deletion target.
19
19
  Prefer `git clean -- <exact prefix>` or tool-native cleanup. A safety prompt that
20
20
  still appears is a hold; agents do not answer it.
@@ -45,7 +45,7 @@ function providerFrom(catalog, provider) {
45
45
  return instance;
46
46
  }
47
47
 
48
- function modelFrom(models, { provider, model: pin, class: modelClass, excluded }) {
48
+ function modelFrom(models, { provider, model: pin, class: modelClass, excluded, effort }) {
49
49
  if (pin && pin !== 'null') {
50
50
  const model = models.find((model) => model.id === pin && !excluded.includes(model.id));
51
51
  if (!model) throw new Error(`missing requested model ${pin}`);
@@ -56,8 +56,9 @@ function modelFrom(models, { provider, model: pin, class: modelClass, excluded }
56
56
  if (['codex', 'claude'].includes(provider)) {
57
57
  throw new Error(`intentional absence for ${provider}: missing model and class`);
58
58
  }
59
- // Launch-by-agent-ID providers bind the first listed ID; exclusions never select a substitute.
60
- const model = models[0];
59
+ // Antigravity encodes effort in its ID; exclusions never select a substitute.
60
+ const model = provider === 'antigravity'
61
+ ? models.find((model) => model.id.endsWith(`-${effort}`)) : models[0];
61
62
  if (!model || excluded.includes(model.id)) throw new Error('missing first listed model');
62
63
  return model;
63
64
  }
@@ -88,14 +89,21 @@ try {
88
89
  const catalog = JSON.parse(readFileSync(path, 'utf8'));
89
90
  const instance = providerFrom(catalog, provider);
90
91
  const model = modelFrom(instance.models, options);
91
- const effortOptions = model.options.filter((option) => effortIds[provider]
92
- ? option.id === effortIds[provider] : ['reasoningEffort', 'effort'].includes(option.id));
93
- if (effortOptions.length !== 1 || !effortOptions[0].options?.some((value) => value.id === effort)
94
- || (provider === 'grok' && effort === 'max')) {
95
- throw new Error(`unsupported effort ${effort} for ${provider}/${model.id}`);
92
+ let effortOption = null;
93
+ if (provider === 'antigravity') {
94
+ if (!model.id.endsWith(`-${effort}`)) {
95
+ throw new Error(`unsupported effort ${effort} for ${provider}/${model.id}`);
96
+ }
97
+ } else {
98
+ const effortOptions = model.options.filter((option) => option.id === effortIds[provider]);
99
+ if (effortOptions.length !== 1 || !effortOptions[0].options?.some((value) => value.id === effort)
100
+ || (provider === 'grok' && effort === 'max')) {
101
+ throw new Error(`unsupported effort ${effort} for ${provider}/${model.id}`);
102
+ }
103
+ effortOption = { id: effortOptions[0].id, value: effort };
96
104
  }
97
105
  console.log(JSON.stringify({ provider, providerInstanceId: instance.providerInstanceId,
98
- model: model.id, effortOption: { id: effortOptions[0].id, value: effort }, source: 'capabilities', path }));
106
+ model: model.id, effortOption, source: 'capabilities', path }));
99
107
  } catch (error) {
100
108
  console.error(`model capabilities resolution hold: ${error.message}`);
101
109
  process.exitCode = 1;
@@ -26,7 +26,9 @@ Set the rung from researched facts; never ask the user to choose it. A change
26
26
  inside one module's existing interface, ownership, data flow, and failure
27
27
  guarantees is Rung 0: no design questions or sketch. Otherwise load the
28
28
  [design lens ladder](../axstack/references/design-lens.md) for Rung 1 or 2
29
- and settle only unresolved areas in its order within the existing budget. Carry a
29
+ and settle only unresolved areas in its order within the existing budget.
30
+ For unresolved Rung 1 or 2 design questions, load
31
+ [Brainstorm](../axstack-brainstorm/SKILL.md) inline. Carry a
30
32
  Rung 1 or 2 sketch in the substantial spec's `Design` section or the returned
31
33
  small-change intent. A design question alone does not make small work
32
34
  substantial; apply routing's existing size reassessment rule.
@@ -65,9 +67,10 @@ substantial; apply routing's existing size reassessment rule.
65
67
  The current chat remains the driver under
66
68
  [Standing contracts](../axstack/references/contracts.md). For each new
67
69
  user round, the driver independently drafts the prioritized frontier and
68
- recommendations, except for an arena-grade question (below), where the driver
69
- writes the brief and rubric but drafts no recommendation until the candidates
70
- and judge verdicts return, so nothing anchors them. Then consult `axstack-advisor-astra` and
70
+ recommendations, except for brainstorm questions, where the driver frames the
71
+ brief and rubric and assesses after the candidates and required judge rounds
72
+ return. Reuse valid brainstorm receipts to replace the adviser consult for
73
+ that question. For other questions, consult `axstack-advisor-astra` and
71
74
  `axstack-advisor-opus` independently, without cross-reading, using the same
72
75
  bounded evidence and question. Each adviser challenges assumptions, edges,
73
76
  omissions, and alternatives; the driver synthesizes disagreements and accepts
@@ -87,60 +90,10 @@ the draft unchanged; it needs no new adviser pair. Changed draft text, a
87
90
  blocking finding, or a high-stakes decision requires fresh receipts on the new
88
91
  revision.
89
92
 
90
- ## Arena for hard-to-reverse design choices
91
-
92
- Critique of one draft anchors every reader to that draft's shape. Rung 2 designs
93
- alone enter the arena: they meet the same test as for an ADR (a meaningful,
94
- hard-to-reverse, non-obvious trade-off: architecture, module boundaries, data
95
- model, migration strategy). Replace the critique round for that question with
96
- an arena. Small or routine questions never enter the arena.
97
-
98
- 1. **Frame.** The driver writes the brief (the artifact, its constraints, the
99
- settled decisions it must respect) and three to six gradeable rubric
100
- criteria. Candidates receive only the brief; the rubric is for judging.
101
- 2. **Fan out.** Produce one candidate per configured family independently from the same brief,
102
- without cross-reading: `axstack-advisor-astra`, `axstack-advisor-opus`,
103
- `axstack-arena-candidate-grok`, and `axstack-arena-candidate-antigravity`.
104
- Each gives a design, rationale, and rejected alternatives. The driver authors no candidate.
105
- 3. **Cross-judge.** After every candidate completes, give round 1 judge
106
- `axstack-arena-judge-opus` the anonymized, relabeled candidates and rubric to
107
- score every candidate per criterion and recommend a base with a reason.
108
- The driver compares its own pick with the Opus verdict. Only if the driver
109
- and the Opus judge disagree on the base, or the user rejects the round-1
110
- synthesis,
111
- run round 2 with fresh sessions: `axstack-escalation-fable` and `axstack-arena-judge-astra`
112
- independently score the same anonymized candidates and rubric. Judges never
113
- author, never cross-read each other, and never average verdicts. After round-2
114
- verdicts return, the driver re-picks in step 4 and re-presents in step 6.
115
- 4. **Pick.** The driver reads every candidate end to end and scores per
116
- criterion, not on holistic feel, then compares with the judge verdicts from
117
- each completed round. Agreement confirms the base. On disagreement, re-read
118
- the rationales and decide with a stated reason; never average verdicts or
119
- fabricate consensus.
120
- 5. **Graft.** Walk the losing candidates once more for the one or two ideas
121
- worth porting and fold them into the base by hand so the result stays
122
- coherent under one mental model. Convergence on the same shape is a strong
123
- agreement signal: adopt the consensus shape, no graft. Wide divergence
124
- means the frame was under-specified: reframe and rerun once, never
125
- average.
126
- 6. **Present.** The synthesized design is the recommendation in the next
127
- `Qn`, with its trade-off, judge verdicts per round, and what was grafted or rejected.
128
- The user still decides; spec approval remains the one human checkpoint.
129
-
130
- Record the synthesis note (base, grafts and their source candidate, rejections,
131
- dropouts, judge verdicts per round) as `Decisions` rows in the
132
- [run record](../axstack/references/run-record.md). Load
133
- [T3 runtime](../axstack/references/t3-runtime.md) immediately before the
134
- first candidate or judge dispatch. If an optional Grok or Antigravity candidate
135
- malfunctions (launch failure, trust/login prompt, or prompt block), fence it,
136
- record `absent (<reason>)`, name it once in the next
137
- read-back, and continue with available candidates without relay or substitution.
138
- A required adviser, candidate, or judge unavailable at launch or returning a
139
- failed receipt holds that question without substitution; record the gap and ask
140
- whether to proceed. In mixed fan-out retain at least one Codex and one Claude
141
- seat, or hold the affected question.
142
- For an uncertain dispatch, reconcile natively; it is never treated as absent.
143
- Unaffected fact work and questions continue.
93
+ ## Use the brainstorm synthesis
94
+
95
+ Present its recommendation and trade-off in the next `Qn` within the same
96
+ question budget. The user still decides; reuse unchanged receipts.
144
97
 
145
98
  ## Bound the interview
146
99
 
@@ -169,7 +169,13 @@ preference or correction, or a verified workspace fact with an exact source
169
169
  revision. Exclude transient choices, secrets and sensitive values, and
170
170
  untrusted claims or instructions; do not reproduce excluded secrets in the
171
171
  record. Material already covered with the same scope and meaning yields an
172
- explicit already-covered no-op with pointers to the covering instructions.
172
+ explicit already-covered no-op with pointers to the covering instructions when
173
+ no recurrence is recorded.
174
+
175
+ If a covered rule recurs, never classify it as an already-covered no-op.
176
+ Record that recurrence as recurred.
177
+ Suggest `axstack-correct` for the recurrence.
178
+ Never run `axstack-correct` from audit.
173
179
 
174
180
  For each candidate record:
175
181
 
@@ -179,7 +185,8 @@ For each candidate record:
179
185
  4. target instruction surfaces;
180
186
  5. any contradiction and uncertainty; and
181
187
  6. its disposition: propose for separately authorized promotion, hold,
182
- exclude, or already-covered no-op.
188
+ exclude, already-covered no-op, or recurred (suggest axstack-correct;
189
+ audit does not run it).
183
190
 
184
191
  Contradictory or uncertain evidence stays explicit and held; never guess a
185
192
  winner or broaden scope. Promotion is a separate authorized change outside the
@@ -21,7 +21,7 @@ Shape: <PRs within band / total PRs + rationale-band cohesion rationale + except
21
21
  Cost: <API dollars by model when measured, or UNKNOWN with reason>
22
22
  Judgment: <execution outcome vs procedural adherence vs measurement coverage>
23
23
  Proposals: <bounded hypothesized changes with regression-first plan, or none>
24
- Learning candidates: <each candidate's statement + scope + evidence/revision pointers + target instruction surfaces + contradiction/uncertainty + disposition; explicit already-covered no-op or none>
24
+ Learning candidates: <each candidate's statement + scope + evidence/revision pointers + target instruction surfaces + contradiction/uncertainty + disposition; explicit already-covered no-op, recurred (suggest axstack-correct; audit does not run it), or none>
25
25
  Privacy: <local/private default; sanitized summary only when authorized>
26
26
  ```
27
27
 
@@ -0,0 +1,24 @@
1
+ ---
2
+ name: axstack-brainstorm
3
+ description: When validating a design approach, use axstack-brainstorm to compare independent candidates and return a synthesis.
4
+ ---
5
+
6
+ # Brainstorm
7
+
8
+ Standalone usage: `/axstack-brainstorm <problem + candidate approach>`.
9
+ Run inline in the current T3 driver thread. Never delegate a coordinator.
10
+ This procedure is report-only. Do not implement, prototype, or grant execution
11
+ or spec approval. Never interview the user or invoke Align.
12
+
13
+ Load [Standing contracts](../axstack/references/contracts.md) before acting.
14
+ Set the rung from researched facts under the
15
+ [design lens](../axstack/references/design-lens.md); Rung 2 meets its ADR test.
16
+ Follow the [arena procedure](references/arena.md), preserving settled decisions
17
+ and naming missing evidence. Candidates challenge the supplied approach as
18
+ well as proposing alternatives.
19
+
20
+ Return `recommend | revise | reject | unresolved` with the
21
+ [sketch](../axstack/references/design-lens.md#sketch), including `Rejected` and
22
+ `Open`, plus base, criterion scores, grafts, dropouts and completed judge rounds.
23
+ Return proposed questions to the caller. The caller owns preferences, next
24
+ steps and any approval; an open blocker stays unresolved.
@@ -0,0 +1,56 @@
1
+ ## Arena
2
+
3
+ Every brainstorm runs a light arena. At Rung 1, the driver scores, picks and
4
+ grafts without a judge round. Only at Rung 2, run the judge rounds below for
5
+ hard-to-reverse choices; they replace the critique round for that question.
6
+
7
+ 1. **Frame.** The driver writes the brief (the artifact, its constraints, the
8
+ settled decisions it must respect) and three to six gradeable rubric
9
+ criteria. Every brief requires a premise check and comparison with the
10
+ smallest change and doing nothing. The rubric uses agent-contributor red
11
+ flags as evidence prompts: seeing only opened files, copying the nearest
12
+ example, taking the shortest path that compiles. Treat them as prompts, not defects. Candidates receive only the brief; the rubric is for judging.
13
+ 2. **Fan out.** Produce one candidate per configured family independently from the same brief,
14
+ without cross-reading: `axstack-advisor-astra`, `axstack-advisor-opus`,
15
+ `axstack-arena-candidate-grok`, and `axstack-arena-candidate-antigravity`.
16
+ Each gives a design, rationale, and rejected alternatives. The driver authors no candidate.
17
+ The driver drafts no recommendation until every candidate returns.
18
+ 3. **Cross-judge (Rung 2 only).** After every candidate completes, give round 1 judge
19
+ `axstack-arena-judge-opus` the anonymized, relabeled candidates and rubric to
20
+ score every candidate per criterion and recommend a base with a reason.
21
+ The driver compares its own pick with the Opus verdict. Only if the driver
22
+ and the Opus judge disagree on the base, or the caller re-invokes with the user's rejection of the round-1
23
+ synthesis,
24
+ run round 2 with fresh sessions: `axstack-escalation-fable` and `axstack-arena-judge-astra`
25
+ independently score the same anonymized candidates and rubric. Judges never
26
+ author, never cross-read each other, and never average verdicts. After round-2
27
+ verdicts return, the driver re-picks in step 4 and re-presents in step 6.
28
+ 4. **Pick.** The driver reads every candidate end to end and scores per
29
+ criterion, not on holistic feel, then compares with the judge verdicts from
30
+ each completed round. Agreement confirms the base. On disagreement, re-read
31
+ the rationales and decide with a stated reason; never average verdicts or
32
+ fabricate consensus.
33
+ 5. **Graft.** Revisit losing candidates once; graft one or two ideas by hand
34
+ into the coherent base under one mental model. Convergence on the same shape is a strong
35
+ agreement signal: adopt the consensus shape, no graft. Wide divergence
36
+ means the frame was under-specified: reframe and rerun once, never
37
+ average.
38
+ 6. **Present.** Return the synthesized design to the caller with its
39
+ trade-off, judge verdicts per round, and what was grafted or rejected.
40
+ The user still decides; spec approval remains the one human checkpoint.
41
+
42
+ Record the synthesis note (base, graft sources, rejections, dropouts, judge
43
+ verdicts per round) as `Decisions` rows in the
44
+ [caller's run record](../../axstack/references/run-record.md) when one exists,
45
+ otherwise include it in the returned verdict. Load
46
+ [T3 runtime](../../axstack/references/t3-runtime.md) immediately before the
47
+ first candidate or judge dispatch. If an optional Grok or Antigravity candidate
48
+ malfunctions (launch failure, trust/login prompt, or prompt block), fence it,
49
+ record `absent (<reason>)`, name it once in the next
50
+ read-back, and continue with available candidates without relay or substitution.
51
+ A required adviser, candidate, or judge needed for this rung unavailable at
52
+ launch or returning a failed receipt holds that question without substitution;
53
+ record the gap and return a proposed question to the caller about whether to proceed. In mixed fan-out retain at least one Codex and one Claude
54
+ seat, or hold the affected question.
55
+ For an uncertain dispatch, reconcile natively; it is never treated as absent.
56
+ Unaffected fact work and questions continue.
@@ -0,0 +1,41 @@
1
+ ---
2
+ name: axstack-correct
3
+ description: When repeated mistakes need evidence and stronger checks, use axstack-correct to report correction proposals.
4
+ disable-model-invocation: true
5
+ ---
6
+
7
+ # Correct
8
+
9
+ Usage: /axstack-correct ["<correction>"] [runs=<ids> | last=<N, default 10>]
10
+
11
+ Load [Standing contracts](../axstack/references/contracts.md) before acting.
12
+
13
+ Run only at the user's request.
14
+ Write a report only.
15
+ Never edit files.
16
+ Never dispatch workers.
17
+ Never merge.
18
+ Never read transcripts.
19
+
20
+ Read [run records](../axstack/references/run-record.md): Decisions, deviations, Learnings, FAILED receipts, audit proposals.
21
+ Read git history bounded by selected runs.
22
+ Read only PR reviews named in those records.
23
+ Record selected runs and git bounds.
24
+ Report inaccessible evidence.
25
+
26
+ A class needs at least two distinct incidents.
27
+ Each incident needs a distinct pointer: run id + attempt or file:line, SHA, or PR or review URL.
28
+ Count an incident echoed in several records once.
29
+ If a pointer is missing, report recurrence UNKNOWN and keep the class open.
30
+
31
+ Try fixes in this order: remove copied pattern > script or helper error naming the fix > brief or receipt template > bun test > prose.
32
+ Label each rule `shipped` or `runtime: advisory`.
33
+ Use `shipped` only when a check fails in CI on the mistake itself.
34
+ When only audit can observe a rule, label it `runtime: advisory`.
35
+ A prose-contract test alone leaves a rule `runtime: advisory`.
36
+ Propose fixes through a small-change intent and [axstack-implement](../axstack-implement/SKILL.md).
37
+
38
+ The `## Enforced rules` table in the target repo's AGENTS.md changes only in the PR that adds its check.
39
+ The first row's PR adds a test that fails when any row's enforcement path disappears.
40
+ If a `shipped` check missed a later recorded recurrence, report `enforcement failed`.
41
+ For that failure, propose relabelling the row `runtime: advisory` in a separately reviewed PR.
@@ -22,6 +22,8 @@ read it before phase 7 and before any fan-out.
22
22
  Redact secrets before showing any command, output, or artifact; build loops
23
23
  against environment variables so credentials never appear in what is shown.
24
24
 
25
+ For performance claims only, load [Performance checklist](../axstack/references/performance-checklist.md).
26
+
25
27
  ## Phases
26
28
 
27
29
  Each phase has an observable completion criterion. Skip one only with a
@@ -230,6 +230,11 @@ For each PR:
230
230
  to step 1.
231
231
  Merge-ready also requires a current diligence `PASS` at that head; diligence
232
232
  `FINDINGS` return to the same author within the review round.
233
+ Diligence `UNKNOWN` records the PR as `held` with the reason in the run record
234
+ and blocks merge-ready pending the driver's recorded disposition.
235
+ Before merge-ready, treat a run-record `Learning` that contradicts a shipped
236
+ rule as a finding on the owning PR and hold merge-ready until the contradiction is resolved.
237
+ An unrelated run-record `Learning` leaves merge-ready eligibility unchanged.
233
238
  A round with reviewer `REQUEST_CHANGES` and/or diligence `FINDINGS` increments
234
239
  `repairs` once and counts once toward the third-round hold.
235
240
 
@@ -27,6 +27,8 @@ Sol pair that fails to launch is fenced, recorded `absent (<reason>)`, and
27
27
  named once in the next read-back, then skipped without relay or substitution.
28
28
  In mixed fan-out retain a Codex and a Claude seat or hold the affected work.
29
29
 
30
+ For performance claims only, load [Performance checklist](../axstack/references/performance-checklist.md).
31
+
30
32
  ## 1. Bound discovery
31
33
 
32
34
  1. Start with the user's named subsystem or pain. Otherwise inspect recent
@@ -33,6 +33,8 @@ When the caller is a bounded review-manager PR job, load
33
33
  carry the required escalation field and every eligible peer PR takes a binding
34
34
  `APPROVE` or `REQUEST_CHANGES` verdict under the automation exception below.
35
35
 
36
+ For performance claims only, load [Performance checklist](../axstack/references/performance-checklist.md).
37
+
36
38
  ## Codebase findings mode
37
39
 
38
40
  Use this manual mode for existing code at a pinned exact source revision and a
@@ -62,6 +62,12 @@ and the lifecycle's [audit skill](../axstack-audit/SKILL.md) hook.
62
62
  session to return plain AGREE. Present one
63
63
  reviewable, identified revision for this checkpoint. Its user approval
64
64
  creates the execution baseline.
65
+ Name the draft revision covered by each adviser receipt at the checkpoint.
66
+ If draft text changed and an adviser receipt covers an older revision, hold
67
+ approval until fresh receipts cover the presented revision.
68
+ Except for high-stakes decisions, a change confined to a `Decisions` row
69
+ reuses adviser receipts only while draft text, evidence, scope, and question
70
+ remain unchanged.
65
71
  Before user approval, dispatch `axstack-diligence` under
66
72
  [Diligence](../axstack/references/diligence.md) to check the draft against
67
73
  the Align decisions for anything dropped, added, or softened.