axstack 0.21.0 → 0.23.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +4 -2
- package/docs/installation.md +10 -6
- package/docs/workflows.md +13 -9
- package/package.json +1 -1
- package/skills/axstack/references/automations.md +9 -0
- package/skills/axstack/references/candidate-publication.md +11 -1
- package/skills/axstack/references/contracts.md +7 -6
- package/skills/axstack/references/design-lens.md +3 -3
- package/skills/axstack/references/diligence.md +10 -1
- package/skills/axstack/references/evidence-archive.md +1 -1
- package/skills/axstack/references/lifecycle.md +3 -1
- package/skills/axstack/references/performance-checklist.md +22 -0
- package/skills/axstack/references/routing.md +9 -3
- package/skills/axstack/references/run-record.md +15 -0
- package/skills/axstack/references/ste-writing.md +22 -0
- package/skills/axstack/references/t3-runtime.md +3 -0
- package/skills/axstack/references/workspace-hygiene.md +1 -1
- package/skills/axstack/scripts/resolve-models.js +17 -9
- package/skills/axstack-align/SKILL.md +11 -58
- package/skills/axstack-audit/SKILL.md +9 -2
- package/skills/axstack-audit/references/record.md +1 -1
- package/skills/axstack-brainstorm/SKILL.md +24 -0
- package/skills/axstack-brainstorm/references/arena.md +56 -0
- package/skills/axstack-correct/SKILL.md +41 -0
- package/skills/axstack-debug/SKILL.md +2 -0
- package/skills/axstack-implement/SKILL.md +5 -0
- package/skills/axstack-improve/SKILL.md +2 -0
- package/skills/axstack-review/SKILL.md +2 -0
- package/skills/axstack-spec/SKILL.md +6 -0
package/README.md
CHANGED
|
@@ -14,7 +14,8 @@ coordination. You can start at the phase you need.
|
|
|
14
14
|
|
|
15
15
|
| Layer | Skill | What it does |
|
|
16
16
|
| --- | --- | --- |
|
|
17
|
-
| Plan | [axstack-align](skills/axstack-align/SKILL.md) | Settle scope through questions, a design lens sketch, and
|
|
17
|
+
| Plan | [axstack-align](skills/axstack-align/SKILL.md) | Settle scope through questions, a design lens sketch, and inline brainstorm validation. |
|
|
18
|
+
| Plan | [axstack-brainstorm](skills/axstack-brainstorm/SKILL.md) | Validate an approach with independent candidates; judges for hard choices. |
|
|
18
19
|
| Plan | [axstack-spec](skills/axstack-spec/SKILL.md) | Write and approve an observable specification. |
|
|
19
20
|
| Plan | [axstack-tickets](skills/axstack-tickets/SKILL.md) | Break approved scope into executable tasks. |
|
|
20
21
|
| Build | [axstack-implement](skills/axstack-implement/SKILL.md) | Build with strict TDD and an author → review → repair loop. |
|
|
@@ -22,6 +23,7 @@ coordination. You can start at the phase you need.
|
|
|
22
23
|
| Verify | [axstack-review](skills/axstack-review/SKILL.md) | Review a PR or bounded codebase at an exact revision. |
|
|
23
24
|
| Verify | [axstack-improve](skills/axstack-improve/SKILL.md) | Find evidenced codebase improvements without editing code. |
|
|
24
25
|
| Verify | [axstack-audit](skills/axstack-audit/SKILL.md) | Measure a run's outcomes and evidence gaps. |
|
|
26
|
+
| Verify | [axstack-correct](skills/axstack-correct/SKILL.md) | Report repeated mistakes and propose stronger checks when invoked by the user. |
|
|
25
27
|
| Operate | [axstack-watch](skills/axstack-watch/SKILL.md) | Observe or maintain an existing PR within its authority. |
|
|
26
28
|
| Operate | [axstack-cleanup](skills/axstack-cleanup/SKILL.md) | Retire eligible completed agent resources. |
|
|
27
29
|
| Operate | [axstack-relay](skills/axstack-relay/SKILL.md) | Send an explicit message or authorized notification. |
|
|
@@ -36,7 +38,7 @@ implementation. Research, explanation, and peer review can start directly.
|
|
|
36
38
|
|
|
37
39
|
| Failure mode | How Axstack responds |
|
|
38
40
|
| --- | --- |
|
|
39
|
-
| Wrong thing built | Align rounds clarify the request; a
|
|
41
|
+
| Wrong thing built | Align rounds clarify the request; every brainstorm runs a light arena across configured families. Use judges only for Rung 2 hard-to-reverse choices. |
|
|
40
42
|
| Nobody really reviewed it | Strict TDD checks behavior first; with the mixed preset, cross-provider review checks the exact revision. |
|
|
41
43
|
| Design rot | The design lens sketches boundaries before a build; Improve surfaces evidenced changes later. |
|
|
42
44
|
| Agents left a mess | T3 makes delegation visible, one writer owns each PR, cleanup stays bounded, and a human merges. |
|
package/docs/installation.md
CHANGED
|
@@ -220,16 +220,20 @@ reconciles their findings.
|
|
|
220
220
|
|
|
221
221
|
The mixed checker and `axstack-research-web-google` have provider
|
|
222
222
|
`antigravity`; mixed `axstack-research-x` has provider `grok`. All three use
|
|
223
|
-
`model: null` with notes authorizing their agent-ID routes; T3 resolves
|
|
224
|
-
model from the first entry for that provider in saved capabilities.
|
|
225
|
-
Antigravity model
|
|
223
|
+
`model: null` with notes authorizing their agent-ID routes; T3 resolves Grok's
|
|
224
|
+
exact model from the first entry for that provider in saved capabilities.
|
|
225
|
+
For Antigravity `model:null` roles, select the first listed model whose ID ends
|
|
226
|
+
with `-<effort>` from saved capabilities and record its exact ID.
|
|
227
|
+
Missing effort-suffix matches hold resolution for Antigravity, including empty
|
|
228
|
+
model catalogs. The single-provider presets configure the checker and keep
|
|
226
229
|
both cross-provider research routes as intentional absences. Their
|
|
227
230
|
unavailable adviser and round-2 seat remain explicit same-provider
|
|
228
231
|
`model: null` roles, which do not make installation unready;
|
|
229
232
|
Align and Spec still hold until both Astra and Opus can return independent
|
|
230
|
-
receipts.
|
|
231
|
-
|
|
232
|
-
|
|
233
|
+
receipts. Every brainstorm runs a light arena across configured families.
|
|
234
|
+
Use judges only for Rung 2 hard-to-reverse choices: round 1 needs Opus; round 2,
|
|
235
|
+
if invoked, needs escalation Fable and Astra; a required seat that is unavailable
|
|
236
|
+
holds that round. The current chat drives on whatever
|
|
233
237
|
model runs it; no preset carries a driver role. Every other missing, invalid, unsupported, or unavailable role value holds only
|
|
234
238
|
the affected work. Codex and Claude class resolution reads the saved T3 capabilities catalog via
|
|
235
239
|
`skills/axstack/scripts/resolve-models.js --provider`; missing or malformed
|
package/docs/workflows.md
CHANGED
|
@@ -18,6 +18,8 @@ and PR shape.
|
|
|
18
18
|
Direct routes need no spec ceremony:
|
|
19
19
|
|
|
20
20
|
- `axstack-research` answers one bounded source-backed question.
|
|
21
|
+
- `axstack-correct` reports repeated mistakes and proposes stronger checks.
|
|
22
|
+
Only the user invokes it.
|
|
21
23
|
- `axstack-explain` separates implemented, intended, tested, live, and unknown
|
|
22
24
|
behavior; complex visuals receive exact-artifact QA where applicable.
|
|
23
25
|
- `axstack-improve` returns a small ranked set of evidenced improvement
|
|
@@ -162,15 +164,17 @@ for explicit ownership transfer.
|
|
|
162
164
|
- `axstack-align` maps facts and dependencies, asks prioritized questions, and
|
|
163
165
|
consults Astra and Opus independently with the same bounded evidence and
|
|
164
166
|
question. It synthesizes disagreements and reuses unchanged receipts. For a
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
|
|
172
|
-
the
|
|
173
|
-
|
|
167
|
+
Rung 1 or 2 design question it loads `axstack-brainstorm` inline and reuses
|
|
168
|
+
its receipt instead of consulting twice; Align owns the interview.
|
|
169
|
+
- `axstack-brainstorm` validates an approach standalone or inline in the driver,
|
|
170
|
+
report-only. Every invocation compares independent Astra, Opus, Grok and
|
|
171
|
+
Antigravity candidates, including premise and smallest-change/do-nothing
|
|
172
|
+
checks. The driver scores, picks and grafts; Rung 1 uses no judges. At Rung 2,
|
|
173
|
+
`axstack-arena-judge-opus` judges round 1; disagreement on the base or caller
|
|
174
|
+
re-invocation with the user's rejection triggers fresh Fable/Astra round 2.
|
|
175
|
+
It returns a verdict, sketch and proposed questions, with no interview,
|
|
176
|
+
prototype or execution approval. Required seats hold; optional dropouts are
|
|
177
|
+
fenced. Fable also serves the bounded triggers in
|
|
174
178
|
[Standing contracts](../skills/axstack/references/contracts.md).
|
|
175
179
|
- `axstack-spec` writes observable acceptance, exclusions, decisions, and one
|
|
176
180
|
user-approved revision baseline. Linear through the executor MCP is the
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "axstack",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.23.0",
|
|
4
4
|
"description": "Axstack installer and setup CLI: installs owned chat skills and role data, configures supported harness settings, and checks T3 Code capabilities.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"claude-code",
|
|
@@ -271,6 +271,15 @@ The orphan sweep covers the run record's repositories plus registered repositori
|
|
|
271
271
|
`t3_thread_organize` settle/archive changes metadata only; exact guarded Git
|
|
272
272
|
worktree removal remains separate. Unknown, active or user-taken-over threads,
|
|
273
273
|
ambiguous publication and failed salvage stay preserved.
|
|
274
|
+
|
|
275
|
+
A review-manager pass may retire a predecessor pass's local `t3code/*` branch
|
|
276
|
+
with exact `git branch -D <branch>` only when ALL hold: the predecessor pass is
|
|
277
|
+
settled and proven run/lane-owned by recorded identity, its descendants and
|
|
278
|
+
evidence are settled, no worktree has the branch checked out, no remote
|
|
279
|
+
counterpart exists, its tip equals the recorded tip, and its tip is an ancestor
|
|
280
|
+
of the verified remote default branch. Otherwise hold; all other `git branch -d`
|
|
281
|
+
rules remain unchanged.
|
|
282
|
+
|
|
274
283
|
Past the authorized storage limit (default 20 retained lane worktrees), disable the schedule
|
|
275
284
|
with `update_scheduled_task` using `enabled:false` and hold.
|
|
276
285
|
Read back the disabled schedule with `list_scheduled_tasks`; uncertainty holds.
|
|
@@ -17,7 +17,16 @@ protection; a mismatch holds publication. If the push outcome is ambiguous,
|
|
|
17
17
|
inspect remote state before retrying.
|
|
18
18
|
|
|
19
19
|
Before reviewer dispatch, read the remote ref back and confirm that it resolves
|
|
20
|
-
to the candidate SHA; also pin the current base.
|
|
20
|
+
to the candidate SHA; also pin the current base.
|
|
21
|
+
After every push, before post-push diligence or merge-ready, compare the PR body's
|
|
22
|
+
stated head SHA with `Confirmed remote SHA` and record `PR body head SHA: <sha or none>`.
|
|
23
|
+
Accept `none` for a PR body without a stated head SHA.
|
|
24
|
+
Accept a stated full head SHA only when it equals `Confirmed remote SHA`.
|
|
25
|
+
Accept a stated abbreviated head SHA only when it is a matching prefix of
|
|
26
|
+
`Confirmed remote SHA` with at least 7 hexadecimal characters.
|
|
27
|
+
Hold on a mismatched head SHA or an abbreviation shorter than 7 characters.
|
|
28
|
+
Leave other SHAs in the PR body unchanged.
|
|
29
|
+
Record:
|
|
21
30
|
|
|
22
31
|
```text
|
|
23
32
|
Candidate: <sha>
|
|
@@ -25,6 +34,7 @@ Base: <sha>
|
|
|
25
34
|
Remote ref: <branch>
|
|
26
35
|
Expected-old remote SHA: <sha | absent>
|
|
27
36
|
Confirmed remote SHA: <sha>
|
|
37
|
+
PR body head SHA: <sha or none>
|
|
28
38
|
PR: <url>
|
|
29
39
|
CI: <run ID or URL and triggered/pending/completed status>
|
|
30
40
|
```
|
|
@@ -2,10 +2,12 @@
|
|
|
2
2
|
|
|
3
3
|
Apply these authority, scope, and model rules before consequential action.
|
|
4
4
|
|
|
5
|
+
For user-facing output, follow [STE-inspired writing](ste-writing.md).
|
|
6
|
+
|
|
5
7
|
## Required lifecycle load
|
|
6
8
|
|
|
7
|
-
Except for `axstack-audit` and `axstack-
|
|
8
|
-
must load and follow [Shared lifecycle](lifecycle.md) before acting. When a
|
|
9
|
+
Except for `axstack-audit`, `axstack-relay`, and `axstack-correct`, every
|
|
10
|
+
independently called phase must load and follow [Shared lifecycle](lifecycle.md) before acting. When a
|
|
9
11
|
substantive run ends or reaches a meaningful checkpoint, apply the lifecycle
|
|
10
12
|
audit hook. The audit phase loads these contracts, writes its assigned record,
|
|
11
13
|
and stops; it never audits itself.
|
|
@@ -55,10 +57,9 @@ profile. Record the driver's provider and model in the run record.
|
|
|
55
57
|
|
|
56
58
|
For Align and Spec, the driver forms an independent assessment first, then
|
|
57
59
|
consults `axstack-advisor-astra` and `axstack-advisor-opus` independently with
|
|
58
|
-
the same bounded evidence and question. The
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
verdicts return. The driver synthesizes disagreements,
|
|
60
|
+
the same bounded evidence and question. The exception covers every brainstorm: the driver frames
|
|
61
|
+
its brief and rubric and assesses only after the candidates and required judge
|
|
62
|
+
rounds return. The driver synthesizes disagreements,
|
|
62
63
|
owns the decision, and the user still approves the spec. Reuse each valid
|
|
63
64
|
unchanged receipt; changed evidence, scope, or question requires a fresh
|
|
64
65
|
receipt. If either adviser is unavailable, Align and Spec hold without model or
|
|
@@ -15,15 +15,15 @@ sketch through the scope identity.
|
|
|
15
15
|
not reclassify work: Rung 1 can stay small. An unsettled material design
|
|
16
16
|
question still makes routing reassess size.
|
|
17
17
|
- **Rung 2 — arena.** A Rung 1 design that also meets the existing ADR test:
|
|
18
|
-
a meaningful, hard-to-reverse, non-obvious trade-off. Use
|
|
19
|
-
|
|
18
|
+
a meaningful, hard-to-reverse, non-obvious trade-off. Use [Brainstorm](../../axstack-brainstorm/SKILL.md)
|
|
19
|
+
with its Rung 2 judge rounds for that question.
|
|
20
20
|
|
|
21
21
|
There is no numeric threshold, file-count gate, or class-count gate.
|
|
22
22
|
|
|
23
23
|
## Questions in order
|
|
24
24
|
|
|
25
25
|
Ask only unresolved areas in numbered `Qn` rounds of one to three. Give each a
|
|
26
|
-
recommendation, reason, and trade-off; use Align's
|
|
26
|
+
recommendation, reason, and trade-off; use Align's inline brainstorm synthesis. Design and
|
|
27
27
|
arena questions share its unchanged budget: 20 normally, a justified extension
|
|
28
28
|
to 35, then opted-in refinement of at most five.
|
|
29
29
|
|
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
Read the [T3 runtime boundary](t3-runtime.md) before dispatch.
|
|
4
4
|
Dispatch `axstack-diligence` with async `delegate_task` in the driver worktree,
|
|
5
5
|
using a pinned brief and private evidence paths.
|
|
6
|
-
It is read-only, never authors or edits, and returns `PASS` or `
|
|
6
|
+
It is read-only, never authors or edits, and returns `PASS`, `FINDINGS`, or `UNKNOWN`
|
|
7
7
|
with locations, observed evidence, and limits. A stale or missing receipt is
|
|
8
8
|
not a pass. Keep its first pass independent of other reviewers and workers.
|
|
9
9
|
|
|
@@ -21,5 +21,14 @@ publication, compare the author receipt with its evidence folder: red/green
|
|
|
21
21
|
logs exist, and counts, SHAs, and paths match. For release preparation, compare
|
|
22
22
|
the release PR body with the merged PRs.
|
|
23
23
|
|
|
24
|
+
Retain the full output of every full-suite run as a named log in the dispatch's
|
|
25
|
+
private evidence folder.
|
|
26
|
+
For an unattributed full-suite failure, report `UNKNOWN` and return the failure to the driver.
|
|
27
|
+
Never report `PASS` for an unattributed full-suite failure.
|
|
28
|
+
The driver records the disposition of each full-suite failure in the run record.
|
|
29
|
+
Report an attributed failure with its retained output normally.
|
|
30
|
+
An observed full-suite failure still fails the suite, including when attributed or reported `UNKNOWN`.
|
|
31
|
+
No reruns are required.
|
|
32
|
+
|
|
24
33
|
`FINDINGS` identifies a mismatch for the driver to resolve at the owning phase;
|
|
25
34
|
it does not edit the artifact or create another review round by itself.
|
|
@@ -118,6 +118,6 @@ After preservation and salvage checks, archive the exact eligible T3 thread
|
|
|
118
118
|
with `t3_thread_organize`, then remove its exact recorded checkout path using
|
|
119
119
|
`git worktree remove <path>` without force. Verify absence with
|
|
120
120
|
`git worktree list --porcelain`; retire only eligible local-only branches with
|
|
121
|
-
`git branch -d` under workspace hygiene. Never use shell recursive deletion or
|
|
121
|
+
`git branch -d` under workspace hygiene. Never use shell recursive deletion to remove a whole worktree or
|
|
122
122
|
treat archive success as ownership, settlement, liveness or cleanup proof.
|
|
123
123
|
Failure or uncertainty preserves the resource.
|
|
@@ -143,6 +143,8 @@ nothing without tested independent review.
|
|
|
143
143
|
|
|
144
144
|
## Close-out
|
|
145
145
|
|
|
146
|
+
Close-out requires the [close-out acceptance table](run-record.md#close-out-acceptance).
|
|
147
|
+
|
|
146
148
|
PRs merge by forge state; close out: (1) settle every T3 worker run and archive eligible threads; (2) compact record with counts and denominators—user
|
|
147
149
|
interventions/deviations from plan/repairs; (3) `axstack-auditor`: settle
|
|
148
150
|
non-zero/requested, else `counts zero`. A base auditor preflight rejection
|
|
@@ -153,7 +155,7 @@ use `axstack-cleanup`, remove the run's own scheduled tasks under
|
|
|
153
155
|
[Workspace hygiene](workspace-hygiene.md), and close selected external-tracker tickets;
|
|
154
156
|
(5) mark the
|
|
155
157
|
[Run record](run-record.md) `Archived`. `Archived`—one each:
|
|
156
|
-
settlement receipt; compact record path; auditor decision plus settlement
|
|
158
|
+
settlement receipt; compact record path; close-out acceptance table; auditor decision plus settlement
|
|
157
159
|
receipt, `counts zero`, or the unlaunchable UNKNOWN archive receipt; scheduled-task,
|
|
158
160
|
release, and ticket receipts; archive timestamp.
|
|
159
161
|
`active`/receipt-incomplete record: close-out pending, never done. One-step
|
|
@@ -0,0 +1,22 @@
|
|
|
1
|
+
# Performance checklist
|
|
2
|
+
|
|
3
|
+
Use this checklist only for performance claims.
|
|
4
|
+
Keep workflow count accounting with axstack-audit.
|
|
5
|
+
A limiter is the resource or code path that bounds performance.
|
|
6
|
+
|
|
7
|
+
1. **Limiter: What bounds the result?**
|
|
8
|
+
Identify the resource or code path that limits the result from profiles or counters captured during a run.
|
|
9
|
+
2. **Tuning: Did each option use suitable settings?**
|
|
10
|
+
Check each option with comparable production settings, versions, and data.
|
|
11
|
+
3. **Limits: Is the result physically plausible?**
|
|
12
|
+
Compare the result with hardware limits and the changed code's share of total elapsed time.
|
|
13
|
+
4. **Errors: Did the run succeed?**
|
|
14
|
+
Verify successful work and correct outputs using error counts from the measurement run.
|
|
15
|
+
5. **Reproducibility: Does the difference repeat?**
|
|
16
|
+
Report the median and range from repeated, alternating runs of each option.
|
|
17
|
+
6. **Relevance: Does the result matter to users?**
|
|
18
|
+
Measure the end-to-end user path with realistic data and concurrency.
|
|
19
|
+
7. **Work happened: Did the timed work finish?**
|
|
20
|
+
Verify that the intended work completed inside the timed region.
|
|
21
|
+
|
|
22
|
+
Ideas paraphrased from [pstack benchmark-checklist](https://github.com/cursor/plugins/blob/e43c7ee26e0038c6c1fa8380dd34ce86ff94cb2a/pstack/skills/benchmark-checklist/SKILL.md) (MIT), using Brendan Gregg's seven benchmark questions.
|
|
@@ -23,9 +23,11 @@ Resolve Codex and Claude classes with
|
|
|
23
23
|
`skills/axstack/scripts/resolve-models.js --provider <provider> --capabilities <path>`
|
|
24
24
|
using saved T3 capabilities JSON; missing or malformed capabilities holds.
|
|
25
25
|
Claude exact IDs come from capabilities, replacing transcript read-back.
|
|
26
|
-
Use an explicit model as given; for `model:null` without a class,
|
|
27
|
-
|
|
28
|
-
|
|
26
|
+
Use an explicit model as given; for `model:null` without a class, grok uses the
|
|
27
|
+
provider's first listed model from saved capabilities and records its exact ID.
|
|
28
|
+
For `model:null` without a class, Antigravity must select the first listed model
|
|
29
|
+
whose ID ends with `-<effort>` from saved capabilities and record its exact ID.
|
|
30
|
+
Missing effort-suffix matches hold resolution for Antigravity.
|
|
29
31
|
A Codex or Claude role with neither model nor class is an intentional absence
|
|
30
32
|
and holds.
|
|
31
33
|
Resume must reuse the saved capabilities and role snapshot with no re-resolution; changes require the user’s explicit decision.
|
|
@@ -50,6 +52,8 @@ step (3) for user routing: no substitution or same-provider review.
|
|
|
50
52
|
|
|
51
53
|
## Direct routes (no spec ceremony)
|
|
52
54
|
|
|
55
|
+
- Validate an approach -> `axstack-brainstorm`: inline, report-only independent
|
|
56
|
+
candidates; light arena always, judges only at Rung 2; return to the caller.
|
|
53
57
|
- Bounded research -> `axstack-research`: verify primary sources and code,
|
|
54
58
|
cite limits, and fan out distinct questions.
|
|
55
59
|
- Understand a system or gap -> `axstack-explain`:
|
|
@@ -61,6 +65,8 @@ step (3) for user routing: no substitution or same-provider review.
|
|
|
61
65
|
off a classified repair (explain: how; debug: what's wrong).
|
|
62
66
|
- Code quality/refactor discovery -> `axstack-improve`: rank bounded
|
|
63
67
|
candidates with evidence; report only, no source edits.
|
|
68
|
+
- Repeated mistakes need evidence and stronger checks -> `axstack-correct`:
|
|
69
|
+
user-invoked, report only.
|
|
64
70
|
- Accepted worker/task/run completion or bounded backlog request -> driver invokes
|
|
65
71
|
`axstack-cleanup` inline; never dispatch it.
|
|
66
72
|
- Preparation completion, watch expiry, resume, or reconciliation -> the
|
|
@@ -119,6 +119,21 @@ by the next owner:
|
|
|
119
119
|
evidence lives. Read it on resume before reconciling; append, never rewrite,
|
|
120
120
|
and keep entries as short as the evidence pointer allows.
|
|
121
121
|
|
|
122
|
+
## Close-out acceptance
|
|
123
|
+
|
|
124
|
+
Use one row per acceptance check, including each clause of a compound check.
|
|
125
|
+
Record in each row a passing evidence pointer or a user-accepted hold with its `Decisions` row.
|
|
126
|
+
If neither passing evidence nor a `Decisions` row with a user-accepted hold exists for an acceptance clause, hold close-out.
|
|
127
|
+
A recorded user-accepted hold in `Decisions` satisfies that clause for close-out; keep the unmet result explicit.
|
|
128
|
+
Record the reason for each driver-elected repair in `Decisions`.
|
|
129
|
+
|
|
130
|
+
```markdown
|
|
131
|
+
| Acceptance check | Result | Evidence or user-accepted hold | Repair election (Decisions row + reason, or none) |
|
|
132
|
+
| --- | --- | --- | --- |
|
|
133
|
+
| <check + clause> | passed | <SHA + check/log pointer> | <Decisions row + reason, or none> |
|
|
134
|
+
| <check + unmet clause> | held (user accepted) | <Decisions row + user acceptance receipt> | <Decisions row + reason, or none> |
|
|
135
|
+
```
|
|
136
|
+
|
|
122
137
|
## Privacy
|
|
123
138
|
|
|
124
139
|
Record concise IDs, SHAs, URLs, status, timestamps, next actions, and evidence
|
|
@@ -0,0 +1,22 @@
|
|
|
1
|
+
# STE-inspired writing
|
|
2
|
+
|
|
3
|
+
STE means Simplified Technical English.
|
|
4
|
+
This reference borrows clarity rules from STE.
|
|
5
|
+
|
|
6
|
+
Follow these rules only for new or materially revised user-facing output.
|
|
7
|
+
User-facing output includes reports, PR bodies, briefs, read-backs, and packaged guidance.
|
|
8
|
+
Do not rewrite existing prose for style.
|
|
9
|
+
Keep exact identifiers unchanged.
|
|
10
|
+
Keep quotes unchanged.
|
|
11
|
+
Keep safety contracts unchanged.
|
|
12
|
+
Keep prompt-byte contracts unchanged.
|
|
13
|
+
|
|
14
|
+
Give each sentence one instruction.
|
|
15
|
+
Keep sentences short.
|
|
16
|
+
Write in active voice.
|
|
17
|
+
If a step has a condition, state the condition before the step.
|
|
18
|
+
Define technical terms before their first use.
|
|
19
|
+
Use one name consistently for each term.
|
|
20
|
+
Use words instead of slashes for "and" or "or".
|
|
21
|
+
|
|
22
|
+
Ideas drawn from [pstack technical-writing](https://github.com/cursor/plugins/blob/e43c7ee26e0038c6c1fa8380dd34ce86ff94cb2a/pstack/skills/technical-writing/SKILL.md#L65-L88) (MIT).
|
|
@@ -32,6 +32,9 @@ A `model:null` role lacking a class must use the first model listed for its
|
|
|
32
32
|
provider in saved capabilities only for grok and antigravity (launch-by-agent-id
|
|
33
33
|
providers); record the exact ID, rather than an unresolved provider default.
|
|
34
34
|
Antigravity must hold when saved capabilities advertise zero models.
|
|
35
|
+
For Antigravity, first-listed selection must use the first model ID ending in `-<effort>` because its model ID encodes effort.
|
|
36
|
+
No matching effort suffix holds resolution for Antigravity.
|
|
37
|
+
For Antigravity, pass no effort option; effort read-back uses the model ID suffix.
|
|
35
38
|
For codex or claude, a role lacking both model and class is an intentional
|
|
36
39
|
absence and must hold; never use a provider default for that role.
|
|
37
40
|
|
|
@@ -14,7 +14,7 @@ Validate its real path, absence of symlinks and ownership before use and cleanup
|
|
|
14
14
|
Evidence files still go to the private `<run>/evidence/<key>/` folder.
|
|
15
15
|
|
|
16
16
|
Every shell deletion targets a literal absolute path or a `${VAR:?}`-guarded expansion, only inside the worker's own evidence folder, `TMPDIR`, or worktree.
|
|
17
|
-
For
|
|
17
|
+
For validated owned scratch, use `rm -r /tmp/<dispatch-key>/scratch` on a literal absolute path inside the evidence folder, `TMPDIR`, or worktree.
|
|
18
18
|
Never use a bare `$VAR`, a glob on a variable, `/`, `HOME`, or a shared root as a deletion target.
|
|
19
19
|
Prefer `git clean -- <exact prefix>` or tool-native cleanup. A safety prompt that
|
|
20
20
|
still appears is a hold; agents do not answer it.
|
|
@@ -45,7 +45,7 @@ function providerFrom(catalog, provider) {
|
|
|
45
45
|
return instance;
|
|
46
46
|
}
|
|
47
47
|
|
|
48
|
-
function modelFrom(models, { provider, model: pin, class: modelClass, excluded }) {
|
|
48
|
+
function modelFrom(models, { provider, model: pin, class: modelClass, excluded, effort }) {
|
|
49
49
|
if (pin && pin !== 'null') {
|
|
50
50
|
const model = models.find((model) => model.id === pin && !excluded.includes(model.id));
|
|
51
51
|
if (!model) throw new Error(`missing requested model ${pin}`);
|
|
@@ -56,8 +56,9 @@ function modelFrom(models, { provider, model: pin, class: modelClass, excluded }
|
|
|
56
56
|
if (['codex', 'claude'].includes(provider)) {
|
|
57
57
|
throw new Error(`intentional absence for ${provider}: missing model and class`);
|
|
58
58
|
}
|
|
59
|
-
//
|
|
60
|
-
const model =
|
|
59
|
+
// Antigravity encodes effort in its ID; exclusions never select a substitute.
|
|
60
|
+
const model = provider === 'antigravity'
|
|
61
|
+
? models.find((model) => model.id.endsWith(`-${effort}`)) : models[0];
|
|
61
62
|
if (!model || excluded.includes(model.id)) throw new Error('missing first listed model');
|
|
62
63
|
return model;
|
|
63
64
|
}
|
|
@@ -88,14 +89,21 @@ try {
|
|
|
88
89
|
const catalog = JSON.parse(readFileSync(path, 'utf8'));
|
|
89
90
|
const instance = providerFrom(catalog, provider);
|
|
90
91
|
const model = modelFrom(instance.models, options);
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
92
|
+
let effortOption = null;
|
|
93
|
+
if (provider === 'antigravity') {
|
|
94
|
+
if (!model.id.endsWith(`-${effort}`)) {
|
|
95
|
+
throw new Error(`unsupported effort ${effort} for ${provider}/${model.id}`);
|
|
96
|
+
}
|
|
97
|
+
} else {
|
|
98
|
+
const effortOptions = model.options.filter((option) => option.id === effortIds[provider]);
|
|
99
|
+
if (effortOptions.length !== 1 || !effortOptions[0].options?.some((value) => value.id === effort)
|
|
100
|
+
|| (provider === 'grok' && effort === 'max')) {
|
|
101
|
+
throw new Error(`unsupported effort ${effort} for ${provider}/${model.id}`);
|
|
102
|
+
}
|
|
103
|
+
effortOption = { id: effortOptions[0].id, value: effort };
|
|
96
104
|
}
|
|
97
105
|
console.log(JSON.stringify({ provider, providerInstanceId: instance.providerInstanceId,
|
|
98
|
-
model: model.id, effortOption
|
|
106
|
+
model: model.id, effortOption, source: 'capabilities', path }));
|
|
99
107
|
} catch (error) {
|
|
100
108
|
console.error(`model capabilities resolution hold: ${error.message}`);
|
|
101
109
|
process.exitCode = 1;
|
|
@@ -26,7 +26,9 @@ Set the rung from researched facts; never ask the user to choose it. A change
|
|
|
26
26
|
inside one module's existing interface, ownership, data flow, and failure
|
|
27
27
|
guarantees is Rung 0: no design questions or sketch. Otherwise load the
|
|
28
28
|
[design lens ladder](../axstack/references/design-lens.md) for Rung 1 or 2
|
|
29
|
-
and settle only unresolved areas in its order within the existing budget.
|
|
29
|
+
and settle only unresolved areas in its order within the existing budget.
|
|
30
|
+
For unresolved Rung 1 or 2 design questions, load
|
|
31
|
+
[Brainstorm](../axstack-brainstorm/SKILL.md) inline. Carry a
|
|
30
32
|
Rung 1 or 2 sketch in the substantial spec's `Design` section or the returned
|
|
31
33
|
small-change intent. A design question alone does not make small work
|
|
32
34
|
substantial; apply routing's existing size reassessment rule.
|
|
@@ -65,9 +67,10 @@ substantial; apply routing's existing size reassessment rule.
|
|
|
65
67
|
The current chat remains the driver under
|
|
66
68
|
[Standing contracts](../axstack/references/contracts.md). For each new
|
|
67
69
|
user round, the driver independently drafts the prioritized frontier and
|
|
68
|
-
recommendations, except for
|
|
69
|
-
|
|
70
|
-
|
|
70
|
+
recommendations, except for brainstorm questions, where the driver frames the
|
|
71
|
+
brief and rubric and assesses after the candidates and required judge rounds
|
|
72
|
+
return. Reuse valid brainstorm receipts to replace the adviser consult for
|
|
73
|
+
that question. For other questions, consult `axstack-advisor-astra` and
|
|
71
74
|
`axstack-advisor-opus` independently, without cross-reading, using the same
|
|
72
75
|
bounded evidence and question. Each adviser challenges assumptions, edges,
|
|
73
76
|
omissions, and alternatives; the driver synthesizes disagreements and accepts
|
|
@@ -87,60 +90,10 @@ the draft unchanged; it needs no new adviser pair. Changed draft text, a
|
|
|
87
90
|
blocking finding, or a high-stakes decision requires fresh receipts on the new
|
|
88
91
|
revision.
|
|
89
92
|
|
|
90
|
-
##
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
hard-to-reverse, non-obvious trade-off: architecture, module boundaries, data
|
|
95
|
-
model, migration strategy). Replace the critique round for that question with
|
|
96
|
-
an arena. Small or routine questions never enter the arena.
|
|
97
|
-
|
|
98
|
-
1. **Frame.** The driver writes the brief (the artifact, its constraints, the
|
|
99
|
-
settled decisions it must respect) and three to six gradeable rubric
|
|
100
|
-
criteria. Candidates receive only the brief; the rubric is for judging.
|
|
101
|
-
2. **Fan out.** Produce one candidate per configured family independently from the same brief,
|
|
102
|
-
without cross-reading: `axstack-advisor-astra`, `axstack-advisor-opus`,
|
|
103
|
-
`axstack-arena-candidate-grok`, and `axstack-arena-candidate-antigravity`.
|
|
104
|
-
Each gives a design, rationale, and rejected alternatives. The driver authors no candidate.
|
|
105
|
-
3. **Cross-judge.** After every candidate completes, give round 1 judge
|
|
106
|
-
`axstack-arena-judge-opus` the anonymized, relabeled candidates and rubric to
|
|
107
|
-
score every candidate per criterion and recommend a base with a reason.
|
|
108
|
-
The driver compares its own pick with the Opus verdict. Only if the driver
|
|
109
|
-
and the Opus judge disagree on the base, or the user rejects the round-1
|
|
110
|
-
synthesis,
|
|
111
|
-
run round 2 with fresh sessions: `axstack-escalation-fable` and `axstack-arena-judge-astra`
|
|
112
|
-
independently score the same anonymized candidates and rubric. Judges never
|
|
113
|
-
author, never cross-read each other, and never average verdicts. After round-2
|
|
114
|
-
verdicts return, the driver re-picks in step 4 and re-presents in step 6.
|
|
115
|
-
4. **Pick.** The driver reads every candidate end to end and scores per
|
|
116
|
-
criterion, not on holistic feel, then compares with the judge verdicts from
|
|
117
|
-
each completed round. Agreement confirms the base. On disagreement, re-read
|
|
118
|
-
the rationales and decide with a stated reason; never average verdicts or
|
|
119
|
-
fabricate consensus.
|
|
120
|
-
5. **Graft.** Walk the losing candidates once more for the one or two ideas
|
|
121
|
-
worth porting and fold them into the base by hand so the result stays
|
|
122
|
-
coherent under one mental model. Convergence on the same shape is a strong
|
|
123
|
-
agreement signal: adopt the consensus shape, no graft. Wide divergence
|
|
124
|
-
means the frame was under-specified: reframe and rerun once, never
|
|
125
|
-
average.
|
|
126
|
-
6. **Present.** The synthesized design is the recommendation in the next
|
|
127
|
-
`Qn`, with its trade-off, judge verdicts per round, and what was grafted or rejected.
|
|
128
|
-
The user still decides; spec approval remains the one human checkpoint.
|
|
129
|
-
|
|
130
|
-
Record the synthesis note (base, grafts and their source candidate, rejections,
|
|
131
|
-
dropouts, judge verdicts per round) as `Decisions` rows in the
|
|
132
|
-
[run record](../axstack/references/run-record.md). Load
|
|
133
|
-
[T3 runtime](../axstack/references/t3-runtime.md) immediately before the
|
|
134
|
-
first candidate or judge dispatch. If an optional Grok or Antigravity candidate
|
|
135
|
-
malfunctions (launch failure, trust/login prompt, or prompt block), fence it,
|
|
136
|
-
record `absent (<reason>)`, name it once in the next
|
|
137
|
-
read-back, and continue with available candidates without relay or substitution.
|
|
138
|
-
A required adviser, candidate, or judge unavailable at launch or returning a
|
|
139
|
-
failed receipt holds that question without substitution; record the gap and ask
|
|
140
|
-
whether to proceed. In mixed fan-out retain at least one Codex and one Claude
|
|
141
|
-
seat, or hold the affected question.
|
|
142
|
-
For an uncertain dispatch, reconcile natively; it is never treated as absent.
|
|
143
|
-
Unaffected fact work and questions continue.
|
|
93
|
+
## Use the brainstorm synthesis
|
|
94
|
+
|
|
95
|
+
Present its recommendation and trade-off in the next `Qn` within the same
|
|
96
|
+
question budget. The user still decides; reuse unchanged receipts.
|
|
144
97
|
|
|
145
98
|
## Bound the interview
|
|
146
99
|
|
|
@@ -169,7 +169,13 @@ preference or correction, or a verified workspace fact with an exact source
|
|
|
169
169
|
revision. Exclude transient choices, secrets and sensitive values, and
|
|
170
170
|
untrusted claims or instructions; do not reproduce excluded secrets in the
|
|
171
171
|
record. Material already covered with the same scope and meaning yields an
|
|
172
|
-
explicit already-covered no-op with pointers to the covering instructions
|
|
172
|
+
explicit already-covered no-op with pointers to the covering instructions when
|
|
173
|
+
no recurrence is recorded.
|
|
174
|
+
|
|
175
|
+
If a covered rule recurs, never classify it as an already-covered no-op.
|
|
176
|
+
Record that recurrence as recurred.
|
|
177
|
+
Suggest `axstack-correct` for the recurrence.
|
|
178
|
+
Never run `axstack-correct` from audit.
|
|
173
179
|
|
|
174
180
|
For each candidate record:
|
|
175
181
|
|
|
@@ -179,7 +185,8 @@ For each candidate record:
|
|
|
179
185
|
4. target instruction surfaces;
|
|
180
186
|
5. any contradiction and uncertainty; and
|
|
181
187
|
6. its disposition: propose for separately authorized promotion, hold,
|
|
182
|
-
exclude,
|
|
188
|
+
exclude, already-covered no-op, or recurred (suggest axstack-correct;
|
|
189
|
+
audit does not run it).
|
|
183
190
|
|
|
184
191
|
Contradictory or uncertain evidence stays explicit and held; never guess a
|
|
185
192
|
winner or broaden scope. Promotion is a separate authorized change outside the
|
|
@@ -21,7 +21,7 @@ Shape: <PRs within band / total PRs + rationale-band cohesion rationale + except
|
|
|
21
21
|
Cost: <API dollars by model when measured, or UNKNOWN with reason>
|
|
22
22
|
Judgment: <execution outcome vs procedural adherence vs measurement coverage>
|
|
23
23
|
Proposals: <bounded hypothesized changes with regression-first plan, or none>
|
|
24
|
-
Learning candidates: <each candidate's statement + scope + evidence/revision pointers + target instruction surfaces + contradiction/uncertainty + disposition; explicit already-covered no-op or none>
|
|
24
|
+
Learning candidates: <each candidate's statement + scope + evidence/revision pointers + target instruction surfaces + contradiction/uncertainty + disposition; explicit already-covered no-op, recurred (suggest axstack-correct; audit does not run it), or none>
|
|
25
25
|
Privacy: <local/private default; sanitized summary only when authorized>
|
|
26
26
|
```
|
|
27
27
|
|
|
@@ -0,0 +1,24 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: axstack-brainstorm
|
|
3
|
+
description: When validating a design approach, use axstack-brainstorm to compare independent candidates and return a synthesis.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Brainstorm
|
|
7
|
+
|
|
8
|
+
Standalone usage: `/axstack-brainstorm <problem + candidate approach>`.
|
|
9
|
+
Run inline in the current T3 driver thread. Never delegate a coordinator.
|
|
10
|
+
This procedure is report-only. Do not implement, prototype, or grant execution
|
|
11
|
+
or spec approval. Never interview the user or invoke Align.
|
|
12
|
+
|
|
13
|
+
Load [Standing contracts](../axstack/references/contracts.md) before acting.
|
|
14
|
+
Set the rung from researched facts under the
|
|
15
|
+
[design lens](../axstack/references/design-lens.md); Rung 2 meets its ADR test.
|
|
16
|
+
Follow the [arena procedure](references/arena.md), preserving settled decisions
|
|
17
|
+
and naming missing evidence. Candidates challenge the supplied approach as
|
|
18
|
+
well as proposing alternatives.
|
|
19
|
+
|
|
20
|
+
Return `recommend | revise | reject | unresolved` with the
|
|
21
|
+
[sketch](../axstack/references/design-lens.md#sketch), including `Rejected` and
|
|
22
|
+
`Open`, plus base, criterion scores, grafts, dropouts and completed judge rounds.
|
|
23
|
+
Return proposed questions to the caller. The caller owns preferences, next
|
|
24
|
+
steps and any approval; an open blocker stays unresolved.
|
|
@@ -0,0 +1,56 @@
|
|
|
1
|
+
## Arena
|
|
2
|
+
|
|
3
|
+
Every brainstorm runs a light arena. At Rung 1, the driver scores, picks and
|
|
4
|
+
grafts without a judge round. Only at Rung 2, run the judge rounds below for
|
|
5
|
+
hard-to-reverse choices; they replace the critique round for that question.
|
|
6
|
+
|
|
7
|
+
1. **Frame.** The driver writes the brief (the artifact, its constraints, the
|
|
8
|
+
settled decisions it must respect) and three to six gradeable rubric
|
|
9
|
+
criteria. Every brief requires a premise check and comparison with the
|
|
10
|
+
smallest change and doing nothing. The rubric uses agent-contributor red
|
|
11
|
+
flags as evidence prompts: seeing only opened files, copying the nearest
|
|
12
|
+
example, taking the shortest path that compiles. Treat them as prompts, not defects. Candidates receive only the brief; the rubric is for judging.
|
|
13
|
+
2. **Fan out.** Produce one candidate per configured family independently from the same brief,
|
|
14
|
+
without cross-reading: `axstack-advisor-astra`, `axstack-advisor-opus`,
|
|
15
|
+
`axstack-arena-candidate-grok`, and `axstack-arena-candidate-antigravity`.
|
|
16
|
+
Each gives a design, rationale, and rejected alternatives. The driver authors no candidate.
|
|
17
|
+
The driver drafts no recommendation until every candidate returns.
|
|
18
|
+
3. **Cross-judge (Rung 2 only).** After every candidate completes, give round 1 judge
|
|
19
|
+
`axstack-arena-judge-opus` the anonymized, relabeled candidates and rubric to
|
|
20
|
+
score every candidate per criterion and recommend a base with a reason.
|
|
21
|
+
The driver compares its own pick with the Opus verdict. Only if the driver
|
|
22
|
+
and the Opus judge disagree on the base, or the caller re-invokes with the user's rejection of the round-1
|
|
23
|
+
synthesis,
|
|
24
|
+
run round 2 with fresh sessions: `axstack-escalation-fable` and `axstack-arena-judge-astra`
|
|
25
|
+
independently score the same anonymized candidates and rubric. Judges never
|
|
26
|
+
author, never cross-read each other, and never average verdicts. After round-2
|
|
27
|
+
verdicts return, the driver re-picks in step 4 and re-presents in step 6.
|
|
28
|
+
4. **Pick.** The driver reads every candidate end to end and scores per
|
|
29
|
+
criterion, not on holistic feel, then compares with the judge verdicts from
|
|
30
|
+
each completed round. Agreement confirms the base. On disagreement, re-read
|
|
31
|
+
the rationales and decide with a stated reason; never average verdicts or
|
|
32
|
+
fabricate consensus.
|
|
33
|
+
5. **Graft.** Revisit losing candidates once; graft one or two ideas by hand
|
|
34
|
+
into the coherent base under one mental model. Convergence on the same shape is a strong
|
|
35
|
+
agreement signal: adopt the consensus shape, no graft. Wide divergence
|
|
36
|
+
means the frame was under-specified: reframe and rerun once, never
|
|
37
|
+
average.
|
|
38
|
+
6. **Present.** Return the synthesized design to the caller with its
|
|
39
|
+
trade-off, judge verdicts per round, and what was grafted or rejected.
|
|
40
|
+
The user still decides; spec approval remains the one human checkpoint.
|
|
41
|
+
|
|
42
|
+
Record the synthesis note (base, graft sources, rejections, dropouts, judge
|
|
43
|
+
verdicts per round) as `Decisions` rows in the
|
|
44
|
+
[caller's run record](../../axstack/references/run-record.md) when one exists,
|
|
45
|
+
otherwise include it in the returned verdict. Load
|
|
46
|
+
[T3 runtime](../../axstack/references/t3-runtime.md) immediately before the
|
|
47
|
+
first candidate or judge dispatch. If an optional Grok or Antigravity candidate
|
|
48
|
+
malfunctions (launch failure, trust/login prompt, or prompt block), fence it,
|
|
49
|
+
record `absent (<reason>)`, name it once in the next
|
|
50
|
+
read-back, and continue with available candidates without relay or substitution.
|
|
51
|
+
A required adviser, candidate, or judge needed for this rung unavailable at
|
|
52
|
+
launch or returning a failed receipt holds that question without substitution;
|
|
53
|
+
record the gap and return a proposed question to the caller about whether to proceed. In mixed fan-out retain at least one Codex and one Claude
|
|
54
|
+
seat, or hold the affected question.
|
|
55
|
+
For an uncertain dispatch, reconcile natively; it is never treated as absent.
|
|
56
|
+
Unaffected fact work and questions continue.
|
|
@@ -0,0 +1,41 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: axstack-correct
|
|
3
|
+
description: When repeated mistakes need evidence and stronger checks, use axstack-correct to report correction proposals.
|
|
4
|
+
disable-model-invocation: true
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Correct
|
|
8
|
+
|
|
9
|
+
Usage: /axstack-correct ["<correction>"] [runs=<ids> | last=<N, default 10>]
|
|
10
|
+
|
|
11
|
+
Load [Standing contracts](../axstack/references/contracts.md) before acting.
|
|
12
|
+
|
|
13
|
+
Run only at the user's request.
|
|
14
|
+
Write a report only.
|
|
15
|
+
Never edit files.
|
|
16
|
+
Never dispatch workers.
|
|
17
|
+
Never merge.
|
|
18
|
+
Never read transcripts.
|
|
19
|
+
|
|
20
|
+
Read [run records](../axstack/references/run-record.md): Decisions, deviations, Learnings, FAILED receipts, audit proposals.
|
|
21
|
+
Read git history bounded by selected runs.
|
|
22
|
+
Read only PR reviews named in those records.
|
|
23
|
+
Record selected runs and git bounds.
|
|
24
|
+
Report inaccessible evidence.
|
|
25
|
+
|
|
26
|
+
A class needs at least two distinct incidents.
|
|
27
|
+
Each incident needs a distinct pointer: run id + attempt or file:line, SHA, or PR or review URL.
|
|
28
|
+
Count an incident echoed in several records once.
|
|
29
|
+
If a pointer is missing, report recurrence UNKNOWN and keep the class open.
|
|
30
|
+
|
|
31
|
+
Try fixes in this order: remove copied pattern > script or helper error naming the fix > brief or receipt template > bun test > prose.
|
|
32
|
+
Label each rule `shipped` or `runtime: advisory`.
|
|
33
|
+
Use `shipped` only when a check fails in CI on the mistake itself.
|
|
34
|
+
When only audit can observe a rule, label it `runtime: advisory`.
|
|
35
|
+
A prose-contract test alone leaves a rule `runtime: advisory`.
|
|
36
|
+
Propose fixes through a small-change intent and [axstack-implement](../axstack-implement/SKILL.md).
|
|
37
|
+
|
|
38
|
+
The `## Enforced rules` table in the target repo's AGENTS.md changes only in the PR that adds its check.
|
|
39
|
+
The first row's PR adds a test that fails when any row's enforcement path disappears.
|
|
40
|
+
If a `shipped` check missed a later recorded recurrence, report `enforcement failed`.
|
|
41
|
+
For that failure, propose relabelling the row `runtime: advisory` in a separately reviewed PR.
|
|
@@ -22,6 +22,8 @@ read it before phase 7 and before any fan-out.
|
|
|
22
22
|
Redact secrets before showing any command, output, or artifact; build loops
|
|
23
23
|
against environment variables so credentials never appear in what is shown.
|
|
24
24
|
|
|
25
|
+
For performance claims only, load [Performance checklist](../axstack/references/performance-checklist.md).
|
|
26
|
+
|
|
25
27
|
## Phases
|
|
26
28
|
|
|
27
29
|
Each phase has an observable completion criterion. Skip one only with a
|
|
@@ -230,6 +230,11 @@ For each PR:
|
|
|
230
230
|
to step 1.
|
|
231
231
|
Merge-ready also requires a current diligence `PASS` at that head; diligence
|
|
232
232
|
`FINDINGS` return to the same author within the review round.
|
|
233
|
+
Diligence `UNKNOWN` records the PR as `held` with the reason in the run record
|
|
234
|
+
and blocks merge-ready pending the driver's recorded disposition.
|
|
235
|
+
Before merge-ready, treat a run-record `Learning` that contradicts a shipped
|
|
236
|
+
rule as a finding on the owning PR and hold merge-ready until the contradiction is resolved.
|
|
237
|
+
An unrelated run-record `Learning` leaves merge-ready eligibility unchanged.
|
|
233
238
|
A round with reviewer `REQUEST_CHANGES` and/or diligence `FINDINGS` increments
|
|
234
239
|
`repairs` once and counts once toward the third-round hold.
|
|
235
240
|
|
|
@@ -27,6 +27,8 @@ Sol pair that fails to launch is fenced, recorded `absent (<reason>)`, and
|
|
|
27
27
|
named once in the next read-back, then skipped without relay or substitution.
|
|
28
28
|
In mixed fan-out retain a Codex and a Claude seat or hold the affected work.
|
|
29
29
|
|
|
30
|
+
For performance claims only, load [Performance checklist](../axstack/references/performance-checklist.md).
|
|
31
|
+
|
|
30
32
|
## 1. Bound discovery
|
|
31
33
|
|
|
32
34
|
1. Start with the user's named subsystem or pain. Otherwise inspect recent
|
|
@@ -33,6 +33,8 @@ When the caller is a bounded review-manager PR job, load
|
|
|
33
33
|
carry the required escalation field and every eligible peer PR takes a binding
|
|
34
34
|
`APPROVE` or `REQUEST_CHANGES` verdict under the automation exception below.
|
|
35
35
|
|
|
36
|
+
For performance claims only, load [Performance checklist](../axstack/references/performance-checklist.md).
|
|
37
|
+
|
|
36
38
|
## Codebase findings mode
|
|
37
39
|
|
|
38
40
|
Use this manual mode for existing code at a pinned exact source revision and a
|
|
@@ -62,6 +62,12 @@ and the lifecycle's [audit skill](../axstack-audit/SKILL.md) hook.
|
|
|
62
62
|
session to return plain AGREE. Present one
|
|
63
63
|
reviewable, identified revision for this checkpoint. Its user approval
|
|
64
64
|
creates the execution baseline.
|
|
65
|
+
Name the draft revision covered by each adviser receipt at the checkpoint.
|
|
66
|
+
If draft text changed and an adviser receipt covers an older revision, hold
|
|
67
|
+
approval until fresh receipts cover the presented revision.
|
|
68
|
+
Except for high-stakes decisions, a change confined to a `Decisions` row
|
|
69
|
+
reuses adviser receipts only while draft text, evidence, scope, and question
|
|
70
|
+
remain unchanged.
|
|
65
71
|
Before user approval, dispatch `axstack-diligence` under
|
|
66
72
|
[Diligence](../axstack/references/diligence.md) to check the draft against
|
|
67
73
|
the Align decisions for anything dropped, added, or softened.
|