axstack 0.26.0 → 0.27.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (33) hide show
  1. package/README.md +5 -2
  2. package/docs/getting-started.md +5 -1
  3. package/docs/installation.md +1 -0
  4. package/docs/workflows.md +18 -8
  5. package/package.json +1 -1
  6. package/profiles/presets/claude-only.json +3 -3
  7. package/profiles/presets/codex-only.json +3 -3
  8. package/profiles/presets/mixed.json +3 -3
  9. package/skills/axstack/references/autopilot.md +36 -5
  10. package/skills/axstack/references/candidate-publication.md +2 -0
  11. package/skills/axstack/references/contracts.md +10 -0
  12. package/skills/axstack/references/lifecycle.md +5 -0
  13. package/skills/axstack/references/perf-loop.md +45 -0
  14. package/skills/axstack/references/role-roster.md +5 -2
  15. package/skills/axstack/references/routing.md +2 -0
  16. package/skills/axstack/references/t3-runtime.md +34 -3
  17. package/skills/axstack/references/ui-verification.md +51 -8
  18. package/skills/axstack/references/workspace-hygiene.md +5 -0
  19. package/skills/axstack-align/SKILL.md +17 -6
  20. package/skills/axstack-audit/SKILL.md +17 -1
  21. package/skills/axstack-debug/SKILL.md +6 -0
  22. package/skills/axstack-diagram/SKILL.md +6 -4
  23. package/skills/axstack-explain/SKILL.md +17 -13
  24. package/skills/axstack-explain/references/inline-pages.md +77 -0
  25. package/skills/axstack-explain/references/visual-qa.md +2 -0
  26. package/skills/axstack-implement/SKILL.md +13 -0
  27. package/skills/axstack-improve/SKILL.md +4 -0
  28. package/skills/axstack-perf/SKILL.md +18 -0
  29. package/skills/axstack-relay/SKILL.md +6 -0
  30. package/skills/axstack-spec/SKILL.md +22 -4
  31. package/skills/axstack-verify/SKILL.md +151 -0
  32. package/skills/axstack-watch/SKILL.md +2 -1
  33. package/skills/axstack-watch/references/watch-runtime.md +5 -2
package/README.md CHANGED
@@ -76,7 +76,9 @@ and [guides](docs/guides.md) for features, reviews, watches, debugging and relea
76
76
  | Plan | [axstack-tickets](skills/axstack-tickets/SKILL.md) | Break approved scope into executable tasks. |
77
77
  | Build | [axstack-implement](skills/axstack-implement/SKILL.md) | Build with strict TDD and an author → review → repair loop. |
78
78
  | Build | [axstack-debug](skills/axstack-debug/SKILL.md) | Diagnose a bug with a failing check and hand off a bounded repair. |
79
+ | Build | [axstack-perf](skills/axstack-perf/SKILL.md) | Route measured performance work to Debug, Improve, or Implement. |
79
80
  | Verify | [axstack-review](skills/axstack-review/SKILL.md) | Review a PR or bounded codebase at an exact revision. |
81
+ | Verify | [axstack-verify](skills/axstack-verify/SKILL.md) | Create, prove and maintain a repository verification skill. |
80
82
  | Verify | [axstack-improve](skills/axstack-improve/SKILL.md) | Find evidenced codebase improvements without editing code. |
81
83
  | Verify | [axstack-audit](skills/axstack-audit/SKILL.md) | Measure a run's outcomes and evidence gaps. |
82
84
  | Verify | [axstack-correct](skills/axstack-correct/SKILL.md) | Report repeated mistakes and propose stronger checks when invoked by the user. |
@@ -85,9 +87,10 @@ and [guides](docs/guides.md) for features, reviews, watches, debugging and relea
85
87
  | Operate | [axstack-relay](skills/axstack-relay/SKILL.md) | Send an explicit message or authorized notification. |
86
88
  | Understand | [axstack-research](skills/axstack-research/SKILL.md) | Answer one bounded question with sources. |
87
89
  | Understand | [axstack-explain](skills/axstack-explain/SKILL.md) | Explain a system and separate known behavior from gaps. |
88
- | Understand | [axstack-diagram](skills/axstack-diagram/SKILL.md) | Draw Mermaid diagrams or verified interactive archify viewers. |
90
+ | Understand | [axstack-diagram](skills/axstack-diagram/SKILL.md) | Draw SVG/CSS in inline pages, Mermaid for chat/GitHub/docs, or requested archify viewers. |
89
91
 
90
- Interactive viewers use [archify](https://github.com/tt-a1i/archify) (MIT).
92
+ Explanations default to checked inline T3 pages above a short reply.
93
+ Use [archify](https://github.com/tt-a1i/archify) (MIT) only on an explicit viewer request.
91
94
 
92
95
  ## How work stays controlled
93
96
 
@@ -24,7 +24,7 @@ axstack check --harness codex
24
24
  ```
25
25
 
26
26
  Installation writes owned skills and instructions, preserving unrelated content,
27
- and fetches the pinned archify tool. Reload the harness's skills in T3 after
27
+ and fetches the pinned archify tool for explicit viewer requests. Reload the harness's skills in T3 after
28
28
  installation. Use the [installation reference](installation.md) for other
29
29
  harnesses, source installs or conflicts.
30
30
 
@@ -43,6 +43,10 @@ reports the owned instruction binding. Missing Chrome is a warning. Any reported
43
43
  gap exits 1; follow [troubleshooting](installation.md#exit-codes-and-troubleshooting) before
44
44
  starting a task.
45
45
 
46
+ Explanations default to checked inline T3 pages above a short reply.
47
+ Use archify only on an explicit viewer request. Spec approval readbacks also
48
+ use inline pages, identifying the revision you approve.
49
+
46
50
  Check does not prove live provider readiness, schedule activation or mobile delivery.
47
51
  Inside the driver thread, capability discovery, provider configuration read-back
48
52
  and actual execution receipts establish their respective runtime boundaries.
@@ -74,6 +74,7 @@ the settings sidecar preserves the value while any other install still owns it.
74
74
  - `--yes` confirms writes under the user's home directory. Tests use temporary
75
75
  homes and fixtures only.
76
76
 
77
+ Explanations default to inline T3 pages; use archify only on an explicit viewer request.
77
78
  Installation sparse-clones archify at the reviewed full SHA in `src/archify-pin.js`
78
79
  into `<tools-dir>/archify-<sha>`. It verifies HEAD before use and keeps the
79
80
  payload's licence and third-party notices. The manifest owns
package/docs/workflows.md CHANGED
@@ -22,16 +22,25 @@ immediately before dispatch or account re-selection.
22
22
 
23
23
  Direct routes need no spec ceremony:
24
24
 
25
+ - [axstack-verify](../skills/axstack-verify/SKILL.md) routes verification-skill
26
+ creation and maintenance through Implement.
27
+ A single-behavior check uses an existing `verify-<app>` skill or UI verification
28
+ delegation. Creation keeps Implement's scope identity requirements.
29
+ - [axstack-perf](../skills/axstack-perf/SKILL.md) routes performance work to
30
+ Debug, Improve, or Implement through the shared performance loop.
25
31
  - `axstack-research` answers one bounded source-backed question.
26
32
  - `axstack-correct` reports repeated mistakes and proposes stronger checks.
27
33
  Only the user invokes it.
28
34
  - `axstack-explain` separates implemented, intended, tested, live, and unknown
29
- behavior; complex visuals receive exact-artifact QA where applicable.
30
- - [axstack-diagram](../skills/axstack-diagram/SKILL.md) selects Mermaid for chat,
31
- GitHub, and docs, or an interactive viewer for required complex visuals.
32
- Explain loads it for every diagram. Viewers use
33
- [archify](https://github.com/tt-a1i/archify) (MIT) with a pinned tool,
34
- a passing finalize receipt, rendered QA, and node-and-edge source review.
35
+ behavior. Explanations default to checked inline T3 pages above a short reply.
36
+ Consequential or complex claims, archify output, or a user request require
37
+ independent explanation review; interactive pages receive a verifier pass.
38
+ - [axstack-diagram](../skills/axstack-diagram/SKILL.md) selects inline SVG or CSS
39
+ inside pages and Mermaid for chat fallback, GitHub, and docs.
40
+ Explain loads it for every diagram. Use
41
+ [archify](https://github.com/tt-a1i/archify) (MIT) only on an explicit viewer request,
42
+ authored by `axstack-explainer`, with a pinned tool, a passing finalize receipt,
43
+ rendered QA, and node-and-edge source review.
35
44
  - `axstack-improve` returns a small ranked set of evidenced improvement
36
45
  candidates without editing code. Its test-audit lens marks every declaration
37
46
  in one owner boundary R/F/C/D, reports reviewed and eligible counts, and routes
@@ -304,7 +313,8 @@ hold pauses the run.
304
313
  After verified publication readback of every own PR from any Axstack phase,
305
314
  the driver arms or joins its chat-run watch in authorized maintain mode.
306
315
  Explicit stop-after-publication and observation-only requests still apply.
307
- Release and install run only under recorded per-run authority.
316
+ Release and install require recorded authority: standing authority from AGENTS.md
317
+ copied into the run's `Release:` and `Authority:`, or explicit per-run authority.
308
318
  Close-out follows their verified receipts and the watch's end.
309
319
 
310
320
  Use `axstack-watch` chat-run mode to watch every PR raised by this run,
@@ -352,7 +362,7 @@ In team mode it clears only an ineligible base, auto-merge turned off, and an op
352
362
  human or bot comment; it never replaces collaborator approval.
353
363
  PRs changing `.github/`, files a workflow step invokes by path, package.json
354
364
  beyond `version` and `files`, lockfiles, test-runner config, branch-protection or
355
- ruleset config, or `CODEOWNERS` are user-merged on the forge.
365
+ ruleset config, `CODEOWNERS`, or `AGENTS.md` are user-merged on the forge.
356
366
  PRs with a non-`clean` revert line are also user-merged on the forge.
357
367
  Promotion, deploying-base, unknown-base, and peer PRs are also user-merged.
358
368
  These categories are excluded from auto-merge.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "axstack",
3
- "version": "0.26.0",
3
+ "version": "0.27.1",
4
4
  "description": "Axstack installer and setup CLI: installs owned chat skills and role data, configures supported harness settings, and checks T3 Code capabilities.",
5
5
  "keywords": [
6
6
  "claude-code",
@@ -142,7 +142,7 @@
142
142
  "provider": "claude",
143
143
  "modeId": "bypassPermissions",
144
144
  "thinkingOptionId": "high",
145
- "notes": "Complex visual explanation author: traces systems, changes, and implementation gaps in requested artifacts and verifies rendered behavior where applicable. T3 resolves Claude classes from saved capabilities; a Claude rejection holds.",
145
+ "notes": "Archify viewer author: authors archify only on an explicit viewer request, with source fidelity. T3 resolves Claude classes from saved capabilities; a Claude rejection holds.",
146
146
  "modelClass": "sonnet"
147
147
  },
148
148
  {
@@ -151,7 +151,7 @@
151
151
  "provider": "claude",
152
152
  "modeId": "bypassPermissions",
153
153
  "thinkingOptionId": "high",
154
- "notes": "Independent visual explanation reviewer: checks the exact artifact for text and source fidelity. The rendered pass belongs to axstack-ui-verifier. Any artifact change invalidates its review. T3 resolves Claude classes from saved capabilities; a Claude rejection holds.",
154
+ "notes": "Independent explanation reviewer: reviews consequential or complex claims, archify output, or on request. Checks the exact artifact for text and source fidelity. The rendered pass belongs to axstack-ui-verifier. Any artifact change invalidates its review. T3 resolves Claude classes from saved capabilities; a Claude rejection holds.",
155
155
  "modelClass": "sonnet"
156
156
  },
157
157
  {
@@ -160,7 +160,7 @@
160
160
  "provider": "claude",
161
161
  "modeId": "bypassPermissions",
162
162
  "thinkingOptionId": "high",
163
- "notes": "Read-only UI verifier: uses T3 preview_* tools against the given build, URL, or artifact; captures screenshots, interactions, accessibility, desktop/mobile, and reduced-motion evidence in the dispatch evidence folder; returns a verdict with evidence paths. Never edits source. T3 resolves Claude classes from saved capabilities; a Claude rejection holds.",
163
+ "notes": "Read-only UI verifier: uses T3 preview_* tools against the given build, URL, or artifact; captures screenshots, interactions, accessibility, desktop/mobile, and reduced-motion evidence in the dispatch evidence folder; returns a verdict with evidence paths. Performs the inline page interaction pass for every interactive page, bound to the exact bytes and SHA-256. Never edits source. T3 resolves Claude classes from saved capabilities; a Claude rejection holds.",
164
164
  "modelClass": "sonnet"
165
165
  },
166
166
  {
@@ -142,7 +142,7 @@
142
142
  "provider": "codex",
143
143
  "modeId": "full-access",
144
144
  "thinkingOptionId": "high",
145
- "notes": "Complex visual explanation author: traces systems, changes, and implementation gaps in requested artifacts and verifies rendered behavior where applicable. Resolve the class from saved T3 capabilities at run start; rejection and unavailable settings hold without substitution.",
145
+ "notes": "Archify viewer author: authors archify only on an explicit viewer request, with source fidelity. Resolve the class from saved T3 capabilities at run start; rejection and unavailable settings hold without substitution.",
146
146
  "modelClass": "sol"
147
147
  },
148
148
  {
@@ -151,7 +151,7 @@
151
151
  "provider": "codex",
152
152
  "modeId": "full-access",
153
153
  "thinkingOptionId": "xhigh",
154
- "notes": "Independent visual explanation reviewer: checks the exact artifact for text and source fidelity. The rendered pass belongs to axstack-ui-verifier. Any artifact change invalidates its review. Resolve the class from saved T3 capabilities at run start; rejection and unavailable settings hold without substitution.",
154
+ "notes": "Independent explanation reviewer: reviews consequential or complex claims, archify output, or on request. Checks the exact artifact for text and source fidelity. The rendered pass belongs to axstack-ui-verifier. Any artifact change invalidates its review. Resolve the class from saved T3 capabilities at run start; rejection and unavailable settings hold without substitution.",
155
155
  "modelClass": "luna"
156
156
  },
157
157
  {
@@ -160,7 +160,7 @@
160
160
  "provider": "codex",
161
161
  "modeId": "full-access",
162
162
  "thinkingOptionId": "medium",
163
- "notes": "Read-only UI verifier: uses T3 preview_* tools against the given build, URL, or artifact; captures screenshots, interactions, accessibility, desktop/mobile, and reduced-motion evidence in the dispatch evidence folder; returns a verdict with evidence paths. Never edits source. Resolve the class from saved T3 capabilities at run start; rejection and unavailable settings hold without substitution.",
163
+ "notes": "Read-only UI verifier: uses T3 preview_* tools against the given build, URL, or artifact; captures screenshots, interactions, accessibility, desktop/mobile, and reduced-motion evidence in the dispatch evidence folder; returns a verdict with evidence paths. Performs the inline page interaction pass for every interactive page, bound to the exact bytes and SHA-256. Never edits source. Resolve the class from saved T3 capabilities at run start; rejection and unavailable settings hold without substitution.",
164
164
  "modelClass": "sol"
165
165
  },
166
166
  {
@@ -142,7 +142,7 @@
142
142
  "provider": "claude",
143
143
  "modeId": "bypassPermissions",
144
144
  "thinkingOptionId": "high",
145
- "notes": "Complex visual explanation author: traces systems, changes, and implementation gaps in requested artifacts and verifies rendered behavior where applicable. T3 resolves Claude classes from saved capabilities; a Claude rejection holds.",
145
+ "notes": "Archify viewer author: authors archify only on an explicit viewer request, with source fidelity. T3 resolves Claude classes from saved capabilities; a Claude rejection holds.",
146
146
  "modelClass": "sonnet"
147
147
  },
148
148
  {
@@ -151,7 +151,7 @@
151
151
  "provider": "codex",
152
152
  "modeId": "full-access",
153
153
  "thinkingOptionId": "xhigh",
154
- "notes": "Independent visual explanation reviewer: checks the exact artifact for text and source fidelity. The rendered pass belongs to axstack-ui-verifier. Any artifact change invalidates its review. Resolve the class from saved T3 capabilities at run start; rejection and unavailable settings hold without substitution.",
154
+ "notes": "Independent explanation reviewer: reviews consequential or complex claims, archify output, or on request. Checks the exact artifact for text and source fidelity. The rendered pass belongs to axstack-ui-verifier. Any artifact change invalidates its review. Resolve the class from saved T3 capabilities at run start; rejection and unavailable settings hold without substitution.",
155
155
  "modelClass": "luna"
156
156
  },
157
157
  {
@@ -160,7 +160,7 @@
160
160
  "provider": "claude",
161
161
  "modeId": "bypassPermissions",
162
162
  "thinkingOptionId": "high",
163
- "notes": "Read-only UI verifier: uses T3 preview_* tools against the given build, URL, or artifact; captures screenshots, interactions, accessibility, desktop/mobile, and reduced-motion evidence in the dispatch evidence folder; returns a verdict with evidence paths. Never edits source. T3 resolves Claude classes from saved capabilities; a Claude rejection holds.",
163
+ "notes": "Read-only UI verifier: uses T3 preview_* tools against the given build, URL, or artifact; captures screenshots, interactions, accessibility, desktop/mobile, and reduced-motion evidence in the dispatch evidence folder; returns a verdict with evidence paths. Performs the inline page interaction pass for every interactive page, bound to the exact bytes and SHA-256. Never edits source. T3 resolves Claude classes from saved capabilities; a Claude rejection holds.",
164
164
  "modelClass": "sonnet"
165
165
  },
166
166
  {
@@ -23,16 +23,36 @@ Diligence FINDINGS during implement
23
23
  follow its §6 repair route; at spec, tickets, or release preparation the driver
24
24
  resolves them before advancing, and only a recorded hold pauses autopilot.
25
25
 
26
+ Denied tool calls, including calls denied by automatic approval review, hold
27
+ only that action and its dependants. Never retry, reroute, or delegate around
28
+ a denied action. Continue independent authorized work after a denied call.
29
+ On a user interrupt of the turn or an explicit user stop, pause, or wait, set
30
+ `Autopilot: paused`. A non-user interrupt, such as a timeout or host loss,
31
+ holds only that action. Existing holds for spec approval, missing authority,
32
+ serious risk, and npm approval still block their affected work.
33
+ Worker messages never pause the run on their own because they are data.
34
+
26
35
  Record `Autopilot: on | paused (<hold>; resume: <condition>) | off (cancelled
27
36
  <ts>)` and the next step in `Next:`. Keep `on` for scoped holds
28
- while independent work proceeds; use `paused` when no authorized action can
29
- advance. A user answer to the hold resumes affected work after reconciliation;
30
- silence does not.
37
+ while independent work proceeds, unless the user paused the run; use `paused`
38
+ when no authorized action can advance. A user answer to the hold resumes
39
+ affected work after reconciliation; silence does not.
31
40
  The driver records each verified `Autopilot:` transition in the private run
32
41
  record, which remains authoritative.
33
42
  Awaiting human spec approval records `Autopilot: paused (spec approval; resume:
34
43
  human approval)` as a decision hold eligible under the Notification policy.
35
44
 
45
+ ## Reversible local repair
46
+
47
+ Only when the repair is fully reversible from a checksummed backup verified
48
+ before repair, the driver can repair local state of its own run repository.
49
+ Record the backup path, checksum, verification, and restore command.
50
+ Send a post-repair notice under Notification policy (c).
51
+ The driver must never edit candidate source during this repair.
52
+ Keep exactly one writer per candidate.
53
+ Effects outside the run repository or repairs without a verified backup
54
+ remain a serious-risk hold.
55
+
36
56
  ## Phase sequence
37
57
 
38
58
  - Small: small-change intent read-back, with Align only when unclear, then
@@ -99,7 +119,13 @@ Align or spec time is a decision hold before release authority is presented.
99
119
  Show the `Release:` line in the spec for human approval at gate 1, or the
100
120
  small-change intent read-back. Copy that decision to `Authority:` in the run
101
121
  record.
102
- This authority is per run and never carries over to another run or repository.
122
+ AGENTS.md can grant standing release and install authority with a trigger and
123
+ named hosts.
124
+ When a run matches that trigger, copy the standing authority into its `Release:`
125
+ line and `Authority:` without a per-run question.
126
+ Standing authority applies only to that repository.
127
+ Without matching standing authority, obtain explicit per-run release and install
128
+ authority before those actions.
103
129
  The small-change intent read-back names the existing Release and host-mutation
104
130
  authority and explicit hosts; silence cannot fill a missing authority or target.
105
131
 
@@ -116,7 +142,7 @@ installed version.
116
142
  An existing version or tag, failed publish, pending approval, uncertain
117
143
  registry result, missing host access, or failed install verification is a
118
144
  resumable hold, never success. Tagging, publishing, installation, and host
119
- mutation require the recorded per-run authority and their existing checks.
145
+ mutation require the run's recorded authority and their existing checks.
120
146
 
121
147
  ## Resume, cancel, and notify
122
148
 
@@ -130,6 +156,11 @@ Cancellation does not cancel a running author run by inference; let it
130
156
  report, then settle that exact attempt under lifecycle guards without new
131
157
  publication.
132
158
 
159
+ On every resume, apply standing merge delegation under watch §5.
160
+ Never narrow standing merge delegation without a user instruction.
161
+ Explicit user restrictions, including `Auto-merge: off`, chat holds, and user
162
+ instructions, always win.
163
+
133
164
  Follow [Provider bindings](t3-runtime.md#preflight-and-binding) for driver account re-selection on start, resume and run-watch wakes.
134
165
 
135
166
  Use the run's recorded Notification policy through `axstack-relay`.
@@ -18,6 +18,8 @@ inspect remote state before retrying.
18
18
 
19
19
  Before reviewer dispatch, read the remote ref back and confirm that it resolves
20
20
  to the candidate SHA; also pin the current base.
21
+ Before reviewer dispatch, refresh the PR body counts and base from the confirmed
22
+ candidate SHA and current base under [PR shape](pr-shape.md).
21
23
  After every push, before post-push diligence or merge-ready, compare the PR body's
22
24
  stated head SHA with `Confirmed remote SHA` and record `PR body head SHA: <sha or none>`.
23
25
  Accept `none` for a PR body without a stated head SHA.
@@ -106,6 +106,16 @@ the evidence, likely impact, options, and needed user decision. Disagreement
106
106
  or silence is not permission. This remains a prompt contract, not a runtime
107
107
  gate.
108
108
 
109
+ For local repository-state recovery, this exception applies.
110
+ Only when the repair is fully reversible from a checksummed backup verified
111
+ before repair, the driver can repair local state of its own run repository.
112
+ Record the backup path, checksum, verification, and restore command.
113
+ Send a post-repair notice under Notification policy (c).
114
+ The driver must never edit candidate source during this repair.
115
+ Keep exactly one writer per candidate.
116
+ Effects outside the run repository or repairs without a verified backup
117
+ remain a serious-risk hold.
118
+
109
119
  ## Authority
110
120
 
111
121
  - The driver owns scope, cross-PR coordination, integration, and every selected
@@ -40,6 +40,11 @@ Preparation completion/watch expiry writes a record. Ordinary resume
40
40
  reconciles it, keeps the current owner, and launches no native handoff.
41
41
  Only an explicit user request to transfer ownership enters this branch.
42
42
 
43
+ A forked thread starts as an observer.
44
+ Before resuming inherited work, the forked thread reconciles the live original
45
+ driver and writers.
46
+ Apply the ownership rules below before any dispatch or write.
47
+
43
48
  1. Reconcile the [Run record](run-record.md) with T3 threads, runs, Git revisions,
44
49
  forge state, pending receipts and scheduled-task expiries; live owners and
45
50
  current attempt identities beat stale state.
@@ -0,0 +1,45 @@
1
+ # Performance loop
2
+
3
+ Keep one writer per candidate.
4
+ The change steps (revert, commit and implement receipt) run only in `axstack-implement`.
5
+
6
+ Follow this ordered loop:
7
+
8
+ 1. **Freeze and prove sensitivity.**
9
+ Freeze the workload, command and environment before the baseline.
10
+ Prove that the measurement harness can detect a change.
11
+ If the harness cannot detect a change, hold the loop.
12
+ 2. **Measure the baseline.**
13
+ Capture the baseline.
14
+ Vet every number with [Performance checklist](performance-checklist.md).
15
+ 3. **Set acceptance checks.**
16
+ Set a target, a noise criterion and a finite attempt budget as acceptance checks in the owning phase.
17
+ The noise criterion is the minimum gain that distinguishes a win from measurement variation.
18
+ 4. **Choose the cheapest hypothesis.**
19
+ Try these mantras in order, cheapest first:
20
+
21
+ - don't do it
22
+ - don't do it again
23
+ - do it less
24
+ - do it later
25
+ - do it when they're not looking
26
+ - do it concurrently
27
+ - do it cheaper
28
+
29
+ Stop when an earlier mantra meets the target.
30
+ 5. **Test one experiment.**
31
+ Verify one change at a time.
32
+ Keep unchanged correctness checks green.
33
+ Keep a change only when its gain exceeds the noise criterion.
34
+ Revert a rejected experiment before the next one.
35
+ Record each rejected experiment with its measured number in the run record and implement receipt.
36
+ 6. **Keep accepted wins.**
37
+ Make one commit per accepted win.
38
+ A changed harness invalidates earlier comparisons.
39
+ 7. **Stop and report.**
40
+ Stop at the target or when the budget is spent.
41
+ Report the baseline, post-change number, delta and artifact path in the run record and implement receipt.
42
+ Report an unmet target as unmet.
43
+
44
+ Ideas paraphrased from pstack's perf-issue and hillclimb
45
+ [playbooks](https://github.com/cursor/plugins/blob/d0ef80d86795816da932a153458c5dbe192d294e/pstack/skills/poteto-mode/playbooks/) (MIT).
@@ -21,10 +21,13 @@
21
21
  `axstack-arena-judge-opus` judges round 1; `axstack-escalation-fable`/`axstack-arena-judge-astra` judge round 2.
22
22
  High-stakes/trigger: fresh [contract](contracts.md) session.
23
23
  `axstack-auditor` audits; `axstack-checker` reports discrepancies.
24
- - `axstack-explainer`/`axstack-explainer-review`: explain/review.
24
+ - `axstack-explainer` authors archify only on an explicit viewer request.
25
+ - `axstack-explainer-review` reviews consequential or complex claims, archify
26
+ output, or on request.
25
27
  - `axstack-diligence`: read-only [diligence checks](diligence.md) for every PR
26
28
  review round and bounded research, spec, ticket, receipt, and release claims.
27
- - `axstack-ui-verifier`: [UI checks](ui-verification.md).
29
+ - `axstack-ui-verifier`: owns the inline page interaction pass under [UI checks](ui-verification.md)
30
+ for every interactive page, bound to the exact bytes and SHA-256.
28
31
  - `axstack-auditor`/`axstack-research-requirements`/
29
32
  `axstack-research-code`/`axstack-research-web`/
30
33
  `axstack-explore-execution`/`axstack-monitor`:
@@ -59,6 +59,8 @@ step (3) for user routing: no substitution or same-provider review.
59
59
 
60
60
  ## Direct routes (no spec ceremony)
61
61
 
62
+ - Route requests to create or maintain a verification skill to `axstack-verify`.
63
+ - Route "make X faster" to `axstack-perf`.
62
64
  - Validate an approach -> `axstack-brainstorm`: inline, report-only independent
63
65
  candidates; light arena always, judges only at Rung 2; return to the caller.
64
66
  - Bounded research -> `axstack-research`: verify primary sources and code,
@@ -130,10 +130,13 @@ configuration; unrelated configurations remain eligible.
130
130
  | Roles | Mechanism and permitted workspace | Completion |
131
131
  |---|---|---|
132
132
  | Advisers, research, read-only explorers, explainers, diligence, checker, auditor, monitor, arena prose candidates and judges, escalation | Must use async `delegate_task` in the driver worktree, title = dispatch key; tracked and untracked files stay untouched; writes only `<run>/evidence/<key>/` | Native notification followed by persisted `task_status` |
133
- | Reviewers (peer/authored), release checks, debug investigators, execution investigators (`axstack-explore-execution`), UI verifier (`axstack-ui-verifier`) | Must use async `delegate_task`, title = dispatch key; driver makes a disposable detached checkout of candidate SHA and pinned base with `git worktree add --detach <run>/checkouts/<key> <sha>` (plus pinned debug patch); brief requires `cd` into it; only disposable probes write there, outputs go to `<run>/evidence/<key>/` | Same delegated terminal checks |
133
+ | Reviewers (peer/authored), release checks, debug investigators, execution investigators (`axstack-explore-execution`), UI verifier candidate checks (`axstack-ui-verifier`) | Must use async `delegate_task`, title = dispatch key; driver makes a disposable detached checkout of candidate SHA and pinned base with `git worktree add --detach <run>/checkouts/<key> <sha>` (plus pinned debug patch); brief requires `cd` into it; only disposable probes write there, outputs go to `<run>/evidence/<key>/` | Same delegated terminal checks |
134
134
  | Author and repairs | Must use `t3_thread_launch` with `{type:worktree, baseRef:<SHA>, branch:<encoded branch>, startFromOrigin:false}` in their own worktree, kept until PR merges or closes | Writer sends a receipt to the driver; driver verifies terminal run and candidate |
135
135
  | Owner | Driver thread in Driver worktree; never writes tracked candidate source or tests; planning artifacts allowed only for repository Markdown; scope, integration, forge mutations and record | No worker launch |
136
136
 
137
+ The UI verifier checkout row applies to candidate checks, while a page without
138
+ a candidate follows [UI verification](ui-verification.md).
139
+
137
140
  The driver must be the sole run-record writer and enforce one writer per
138
141
  candidate; it never writes tracked candidate source or tests or repairs an author's source.
139
142
  The driver may write planning artifacts (spec, ticket map) in its own worktree when the selected store is repository Markdown.
@@ -143,6 +146,8 @@ The current chat/driver has no role row in any preset.
143
146
  The dispatch key must be `<run>:<role>:<task>:a<n>`, recorded before launch and
144
147
  used as the exact whole T3 title. Substring matches do not establish identity.
145
148
  Each dispatch binds the approved spec or small-change intent, brief, authority, role snapshot, base and candidate to its dispatch key.
149
+ Every dispatch brief must state that dispatched roles never call `html_render`.
150
+ Dispatched roles may call `html_preview` within [UI verification](ui-verification.md).
146
151
 
147
152
  The branch must be `axstack/<run>/<role>/<task>-a<n>`; lowercase each segment
148
153
  and replace every `[^a-z0-9-]` character with `-`. Keep the dispatch key in its
@@ -150,6 +155,10 @@ original form and record both; normalization is never an identity substitute.
150
155
 
151
156
  `baseRef` must always be a commit SHA, never a branch name; T3 renames its
152
157
  `t3code/*` branches. Pin base and candidate before dispatch.
158
+ Fetch the pinned base commit into the launch repository before a writer launch.
159
+ Before a writer launch, verify locally that the pinned base SHA resolves to a
160
+ commit with that exact SHA.
161
+ If fetch or verification fails, hold that launch.
153
162
 
154
163
  Workers must finish with exactly one final marker: `AXSTACK-DONE key=… head=…
155
164
  report=…`, `AXSTACK-FAILED key=… head=… report=…`, or `AXSTACK-QUESTION key=… q=…`.
@@ -171,6 +180,9 @@ Launched writer completion must require terminal `t3_thread_wait` on that run,
171
180
  then candidate checks: non-empty diff, clean tree and named red/green logs.
172
181
  A receipt message alone counts only as progress; it can precede terminal state.
173
182
 
183
+ When a completion arrives, process it in the same driver turn after the required
184
+ terminal checks, delegated `hasPendingChildRuns:false`, and writer candidate checks.
185
+
174
186
  Completion must match the current attempt key and candidate SHA. An older
175
187
  attempt never completes a newer one; stale or duplicate receipts remain
176
188
  evidence, deduplicated by runtime identity. Process the whole delivery before
@@ -234,6 +246,22 @@ Before finishing PR work, including watch end or Close-out, the driver calls
234
246
  `list_thread_pull_requests` and links every missing run PR.
235
247
  Never link unrelated PRs.
236
248
 
249
+ After verified publication readback of an own PR, for every layer of a `gh stack`,
250
+ the driver also calls `t3_thread_update` with `action: link_pull_request`,
251
+ `pullRequest` containing its number, repository and full PR URL, and `threadId`
252
+ set to the author thread, then reads back the PR link with
253
+ `list_thread_pull_requests` for that same author thread.
254
+ A missing or failed author-link readback never triggers author repair: it holds
255
+ only that PR's merge-ready and requires retrying the link.
256
+
257
+ When the forge confirms a run PR merged or closed, in the same driver turn the
258
+ driver settles its author thread with `t3_thread_organize` using `action: settle`
259
+ and reads back `settled: true` with `t3_thread_read` for that thread.
260
+ Apply the existing [Settlement](workspace-hygiene.md#settlement) guards:
261
+ terminal run evidence is required and user-taken-over threads stay untouched.
262
+ A missing or failed settle readback holds watch end and Close-out only for
263
+ that author thread until resolved.
264
+
237
265
  For every watched own PR in chat-run or standalone adopted maintenance, when
238
266
  available the driver calls `watch_pull_request` for T3 wakes on check completion,
239
267
  new comments from others, or branch conflicts, then ends the turn.
@@ -245,8 +273,11 @@ Route native PR wake events through watch §4 and the unchanged §5 readiness pr
245
273
 
246
274
  If `watch_pull_request` is unavailable, fall back to the bound 5-minute schedule
247
275
  and `scripts/pr-digest.js` without a hold.
248
- Keep the schedule cadence unchanged while a native PR watch is active,
249
- including the existing 7-day quiet relaxation.
276
+ When the run waits only on a user decision or a human-only step, including a
277
+ user merge, with no unsettled worker and no PR needing watch events, change
278
+ the run watch cadence to 60 minutes at once.
279
+ When work restarts, restore the normal 5-minute cadence.
280
+ Keep native PR watches.
250
281
  The schedule still reconciles author and review tasks, readiness, and release.
251
282
  When a PR merges or closes, or its watch is torn down, call `unwatch_pull_request`
252
283
  and keep its link.
@@ -1,25 +1,68 @@
1
1
  # UI verification
2
2
 
3
- Every Playwright, browser, or rendered-UI check, including a "confirm it in the
4
- browser" step, goes through async `delegate_task` to `axstack-ui-verifier` from the
5
- run's role snapshot. Give it the exact build, URL, or artifact and the private
3
+ Every Playwright, browser, or rendered-UI check outside the three named author
4
+ exceptions, including a "confirm it in the browser" step, goes through async
5
+ `delegate_task` to `axstack-ui-verifier` from the run's role snapshot. Give it the exact build, URL, or artifact and the private
6
6
  dispatch's evidence folder. The verifier is read-only: it never edits source.
7
7
  The PR writer remains the sole writer.
8
8
 
9
9
  Before using a PR preview URL, the UI verifier must read [PR previews](preview.md).
10
10
 
11
- The sole author browser exception is archify `finalize` as a headless build gate.
11
+ The first named author browser exception is archify `finalize` as a headless build gate.
12
12
  Keep finalize outputs only in the private evidence folder.
13
13
  Never replace the verifier's rendered pass with finalize.
14
- All other browser checks remain delegated to `axstack-ui-verifier`.
15
14
 
16
- Read the [T3 runtime boundary](t3-runtime.md) before dispatch and use its
17
- driver-made disposable detached checkout at the pinned candidate SHA.
15
+ The second named author browser exception is the driver's two `html_preview`
16
+ runs at 728px dark and 360px light.
17
+ Scan the exact bytes of every page for remote src, srcset, and href on loaded
18
+ resources, `@import`, `url()`, `fetch`, XHR, WebSocket, EventSource, and meta refresh.
19
+ Reject scan matches except local url(#id) references such as SVG markers and gradients.
20
+ Require correct layout, theme, zero console errors, empty `missingImages`,
21
+ height, network, and word count.
22
+ Never replace the verifier's interaction pass with `html_preview`.
23
+ Follow [Inline pages](../../axstack-explain/references/inline-pages.md) for the page checks.
24
+
25
+ The third named author browser exception permits dispatched roles to use
26
+ `html_preview` as a non-publishing self-check.
27
+ Dispatched-role self-checks never replace the driver's two previews or the `axstack-ui-verifier` pass.
28
+ Dispatched roles never call `html_render`.
29
+ All other browser checks outside the three named exceptions remain delegated to `axstack-ui-verifier`.
30
+
31
+ Read the [T3 runtime boundary](t3-runtime.md) before dispatch; for candidate
32
+ checks use its driver-made disposable detached checkout at the pinned candidate SHA.
18
33
  The verifier uses T3 `preview_*` tools for rendered checks.
19
- Browser and visual checks must run in the delegated `axstack-ui-verifier` in its own detached checkout; outputs go to its private evidence folder, never the driver worktree.
34
+ Candidate browser and visual checks must run in the delegated `axstack-ui-verifier` in its own detached checkout; outputs go to its private evidence folder, never the driver worktree.
35
+
36
+ For every interactive page, require a real-browser `axstack-ui-verifier` interaction pass.
37
+ Give the evidence path and SHA-256 of the exact bytes in the verifier brief.
38
+ The verifier serves the exact bytes on loopback.
39
+ Open the page with `preview_open` using `open:false`.
40
+ Check clicks, keyboard, focus order, screen-reader names, expanded states, and reduced motion.
41
+ Require that `performance.getEntriesByType('resource')` is empty.
42
+ The detached-checkout rule does not apply to a page without a candidate.
20
43
 
21
44
  Ask for screenshots and observed interactions, accessibility, desktop and
22
45
  mobile layouts, and reduced-motion behavior where relevant. The verifier
23
46
  returns a verdict with evidence paths and names checks it could not run.
24
47
  Keep the verdict tied to the exact artifact or revision; changed bytes need a
25
48
  fresh rendered pass.
49
+
50
+ Use a present `verify-<app>` skill for affected behavior.
51
+ If a PR changes a mapped feature, its author updates the feature file in the same PR.
52
+ For a missing skill, add "suggest `axstack-verify create`" to the receipt's `Unverified:` line.
53
+ Never generate a verification skill automatically.
54
+ For a broken or stale recipe, report the gap.
55
+ Never let a broken or stale recipe waive existing acceptance obligations.
56
+ Keep browser recipes tool-neutral for the verifier.
57
+
58
+ ## Proof standards
59
+
60
+ Apply these standards to the verifier's candidate checks and inline-page interaction pass.
61
+
62
+ Drive the real user path.
63
+ Capture the action, the resulting state and its side effects.
64
+ For a dry-run claim, verify what it skips by observing files, network calls or Git refs.
65
+ Keep evidence in the private evidence folder after cleanup.
66
+
67
+ Ideas paraphrased from pstack's
68
+ [create-verification-skill](https://github.com/cursor/plugins/blob/d0ef80d86795816da932a153458c5dbe192d294e/pstack/skills/create-verification-skill/SKILL.md) (MIT).
@@ -13,6 +13,9 @@ Before use, commands must scope `TMPDIR` to an owned 0700 directory under the sy
13
13
  Validate its real path, absence of symlinks and ownership before use and cleanup; remove it afterwards by literal absolute path.
14
14
  Evidence files still go to the private `<run>/evidence/<key>/` folder.
15
15
 
16
+ Run each cleanup as a single command.
17
+ Never use `rm -f` or chain cleanup commands.
18
+
16
19
  Every shell deletion targets a literal absolute path or a `${VAR:?}`-guarded expansion, only inside the worker's own evidence folder, `TMPDIR`, or worktree.
17
20
  For validated owned scratch, use `rm -r /tmp/<dispatch-key>/scratch` on a literal absolute path inside the evidence folder, `TMPDIR`, or worktree.
18
21
  Never use a bare `$VAR`, a glob on a variable, `/`, `HOME`, or a shared root as a deletion target.
@@ -35,6 +38,8 @@ or launched run completion under the runtime contract before settlement.
35
38
  After the driver accepts a launched writer's or delegated task's completion from terminal run evidence plus a verified receipt or candidate check, it settles the matching thread with `t3_thread_organize` as metadata only.
36
39
  Settling never removes, archives or abandons the thread or worktree.
37
40
  Retain the author worktree and thread unarchived until the PR merges or closes, with repairs returning to the same author.
41
+ On forge-confirmed merge or closure, follow [Native PR links and watches](t3-runtime.md#native-pr-links-and-watches)
42
+ for same-turn author settlement and its required readback.
38
43
  A repair turn automatically un-settles the thread, making it visible while working.
39
44
  The driver settles the thread again after the next accepted repair completion.
40
45
  Threads with unaccepted completion, FAILED or QUESTION markers, held, failed or interrupted runs, or anything needing the user stay unsettled and visible.
@@ -52,7 +52,10 @@ substantial; apply routing's existing size reassessment rule.
52
52
  scenarios without using a fixed questionnaire or padding the interview. An
53
53
  empty ready frontier means completion only when no material choice remains;
54
54
  otherwise report the blocking research and continue safe fact work.
55
- 3. **Ask a focused round.** Present one to three independent questions; present
55
+ 3. **Ask a focused round.** Ask only unresolved, consequential preferences or
56
+ authority. Apply preferences fixed by user instructions, `AGENTS.md`, or a
57
+ prior decision within its scope, and list them as defaulted in the read-back.
58
+ Defaults never grant authority. Present one to three independent questions; present
56
59
  one alone when it is complex or governs dependent branches. Number questions
57
60
  cumulatively as `Q1`, `Q2`, and so on. Recommend a choice for each with a
58
61
  short reason and trade-off, then wait for the user's answers and recompute
@@ -86,9 +89,16 @@ and user-resolved choices for `axstack-spec`. If either adviser is unavailable,
86
89
  hold Align; safe fact work may continue without substitution.
87
90
 
88
91
  An optional adviser note may be deferred or rejected in a `Decisions` row with
89
- the draft unchanged; it needs no new adviser pair. Changed draft text, a
90
- blocking finding, or a high-stakes decision requires fresh receipts on the new
91
- revision.
92
+ the draft unchanged; it needs no new adviser pair.
93
+ By default, changed draft text requires fresh receipts on the new revision.
94
+ The only exception to this default is delta confirmation, available only after
95
+ each configured adviser seat confirms the driver's meaning-preservation claim.
96
+ If the driver claims an edit preserves criterion, scope, decision and
97
+ instruction meaning, each configured adviser seat confirms that claim on the
98
+ delta, naming its prior receipt, the delta and the new revision.
99
+ If a seat disagrees, a blocking finding exists, a high-stakes decision arises,
100
+ or meaning changes, obtain full fresh review on the new revision.
101
+ An unavailable adviser seat holds the affected phase.
92
102
 
93
103
  ## Use the brainstorm synthesis
94
104
 
@@ -105,8 +115,9 @@ remaining budget for the highest-value branches and never exceed 35 questions
105
115
  in the initial pass. Stop earlier as soon as no unresolved material choice
106
116
  remains; 20 is not a quota.
107
117
 
108
- At completion or the 35-question cap, read back the result and ask whether the
109
- user wants deeper refinement. Ask no further interview or refinement questions
118
+ At completion or the 35-question cap, read back the result.
119
+ Never make a deeper-refinement offer a completion requirement.
120
+ Ask no further interview or refinement questions
110
121
  without opt-in; this does not replace required spec approval or a clarification
111
122
  prompt when a configured model is unavailable. An opted-in refinement names
112
123
  one area and a separate finite budget of at most five questions; it preserves
@@ -150,7 +150,7 @@ Each proposal names:
150
150
 
151
151
  1. the observed failure or inefficiency;
152
152
  2. the hypothesized root cause, with evidence and counterevidence;
153
- 3. one bounded hypothesized skill change;
153
+ 3. one bounded hypothesized skill or environment change;
154
154
  4. a regression scenario first, followed by an unchanged holdout evaluation;
155
155
  5. a cost and quality comparison when those values were measured; and
156
156
  6. the authorized delivery path: the auditor suggests, the driver arranges an
@@ -160,6 +160,22 @@ Keep evaluation data, candidate changes, and validation separate. This is an
160
160
  original Axstack workflow with no outside dependency or extra framework to
161
161
  install.
162
162
 
163
+ The environment lens is optional for bounded proposals.
164
+ Consider navigation pointers, automated checks, coding-standard placement for review, steering-file bloat and no-op instructions, tool economy, and information access.
165
+ For steering-file trims, follow section 6's AGENTS.md and CLAUDE.md parity rule.
166
+ Route repeated mistakes to axstack-correct as section 6 proposals.
167
+
168
+ Before proposing a new check, report whether an existing check is unwired or broken.
169
+ For a mechanical rule, prefer a deterministic check over a prose rule.
170
+
171
+ For this lens, read only the evidence section 2 permits plus read-only AGENTS.md and CLAUDE.md, CI configuration and package scripts at the audited revision.
172
+ Read no transcripts for the environment lens.
173
+ For a finding without a pointer, report UNKNOWN.
174
+ The environment lens makes no automatic edit.
175
+ For the environment lens, keep the existing delivery path in field 6.
176
+
177
+ Environment ideas paraphrased from Matt Pocock's [retro](https://github.com/mattpocock/skills/blob/6fd947921b935b7e1e69293a200400f0fdd5c15f/skills/engineering/retro/SKILL.md) (MIT).
178
+
163
179
  Omit any proposal that is not testable, does not preserve unchanged
164
180
  expectations, or would grant the auditor implementation or activation
165
181
  authority.
@@ -24,6 +24,12 @@ against environment variables so credentials never appear in what is shown.
24
24
 
25
25
  For performance claims only, load [Performance checklist](../axstack/references/performance-checklist.md).
26
26
 
27
+ For performance work only, load [Performance loop](../axstack/references/perf-loop.md).
28
+ For a performance regression, freeze the workload, command and environment before the baseline.
29
+ For a performance regression, rank hypotheses in mantra order.
30
+ For a performance regression, hand off the repair to `axstack-implement` with real red-to-green evidence.
31
+ For performance work, hand off without committing.
32
+
27
33
  ## Phases
28
34
 
29
35
  Each phase has an observable completion criterion. Skip one only with a
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: axstack-diagram
3
- description: When an explanation needs a diagram, use axstack-diagram to choose Mermaid or a pinned archify viewer with source fidelity and rendered QA.
3
+ description: When an explanation needs a diagram, use axstack-diagram for inline SVG or CSS in pages, Mermaid in chat, GitHub, and docs, or archify only on an explicit viewer request.
4
4
  ---
5
5
 
6
6
  # Diagram
@@ -9,9 +9,11 @@ Usage: `/axstack-diagram <question + destination>` or load this skill from expla
9
9
  Load [Standing contracts](../axstack/references/contracts.md) before acting.
10
10
  Follow [Fidelity](references/fidelity.md) for every format.
11
11
 
12
- Use archify HTML for an explicit viewer request or explain's complex-visual path.
13
- For the viewer, follow [Archify](references/archify.md).
14
- Otherwise use Mermaid for chat, GitHub, or docs.
12
+ Use archify HTML only on an explicit viewer request.
13
+ The configured `axstack-explainer` authors the viewer; follow [Archify](references/archify.md).
14
+ Use inline SVG or CSS inside inline pages.
15
+ Never use Mermaid in inline pages.
16
+ Use Mermaid for chat fallback, GitHub, and docs.
15
17
  Include `accTitle` and `accDescr` in each Mermaid view.
16
18
  Never use theme init in Mermaid.
17
19
  Split large Mermaid diagrams into views.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: axstack-explain
3
- description: When understanding a system, change, or implementation gap, use axstack-explain to show how it works and what exists, is missing, or remains unverified.
3
+ description: When the user asks how, why, explain, or show, says they do not understand, a spec is being finalized, or you judge they need to understand something, use axstack-explain to show how it works and what exists, is missing, or remains unverified.
4
4
  ---
5
5
 
6
6
  # Explain
@@ -71,24 +71,28 @@ Diagram never calls explain.
71
71
  Each card has at most 40 words and cites its source, as defined in
72
72
  [Archify](../axstack-diagram/references/archify.md).
73
73
  Count node labels, headings, captions, and all non-card text.
74
- 2. For a simple request, answer concisely in the current chat. Use a compact
75
- diagram when useful. This needs no mandatory agent or intermediate artifact.
76
- For simple chat answers, never dispatch an agent when applying its Mermaid rules inline.
77
- 3. For a complex visual, use the configured `axstack-explainer` role to create
78
- self-contained HTML through the archify path by default.
79
- For a complex visual, honor an explicitly requested artifact format.
80
- An explicit user theme wins; otherwise use the dark default.
74
+ 2. Use chat only for an explicit chat request, an answer without a Q2 trigger,
75
+ or inline-page fallback. Q2 triggers are the conditions in the description.
76
+ Use a compact diagram when useful.
77
+ For chat answers, never dispatch an agent when applying its Mermaid rules inline.
78
+ 3. For a Q2 trigger, create an inline page by default.
79
+ Honor an explicitly requested artifact format.
80
+ Use archify only on an explicit viewer request, through the configured
81
+ `axstack-explainer`. The reply adds only what the page omits.
81
82
  4. Profile IDs are presets, not availability proof. Before dispatch, follow the
82
83
  launch sequence and preserve the configured model, mode, and effort. Report
83
84
  an unavailable route; never substitute a model.
84
85
 
85
86
  ## 3. Verify and deliver
86
87
 
87
- 1. Any HTML explanation requires the full [visual QA
88
- checklist](references/visual-qa.md): actual desktop and mobile rendering,
89
- interaction, accessibility, and reduced-motion checks where relevant.
90
- 2. For archify output, require the configured independent `axstack-explainer-review`;
91
- for other artifacts, use it when warranted. Bind it to the exact artifact identity.
88
+ 1. For inline pages, follow [Inline pages](references/inline-pages.md).
89
+ For archify and every other explicitly requested HTML or visual artifact
90
+ outside inline T3 pages, follow [Visual QA](references/visual-qa.md).
91
+ 2. Require independent `axstack-explainer-review` for consequential or complex
92
+ claims, archify output, or on request.
93
+ Consequential claims include gap or missing-claim reports, blast-radius
94
+ safety facts, and spec readbacks.
95
+ Bind it to the exact artifact identity.
92
96
  Any byte change invalidates
93
97
  that review and requires a fresh check. In `claude-only`, separate Sonnet
94
98
  author high and reviewer high sessions are allowed for explanations as
@@ -0,0 +1,77 @@
1
+ # Inline pages
2
+
3
+ Use this reference for Explain's inline T3 answer path.
4
+ The driver publishes the page in its own thread.
5
+
6
+ ## Shape and theme
7
+
8
+ Use T3 theme variables and follow the live theme.
9
+ Only an explicit user theme request overrides the page theme.
10
+ Use `--background`, `--foreground`, `--muted-foreground`, `--ring`,
11
+ `--font-sans`, and `--font-mono` for their named purposes.
12
+ The dark default stays archify-only.
13
+
14
+ Use fluid width, with zero outer horizontal padding.
15
+ Never use an outer frame, banner, or viewport heights.
16
+ Let content set the page height. Use fixed chart heights.
17
+ Set frame `height` to the larger contentHeight of the two previews, capped at 2000.
18
+
19
+ Choose the page shape from its content rather than a fixed template.
20
+ Map stages to a connected diagram, sequence to a timeline, comparison to side
21
+ by side or table, and separate equal items to a list.
22
+ Prefer diagrams and timelines. Avoid grids of bordered boxes.
23
+ Keep borders minimal and use color only where it has meaning.
24
+
25
+ ## Page rules
26
+
27
+ Write one self-contained HTML document with inline styles and scripts.
28
+ Never load network resources, including CDN scripts, fonts, images, or Mermaid.
29
+ Use images as `data:` URIs. Never include absolute local paths in page bytes.
30
+ Allow plain `<a href>` links, with GitHub URLs or `path@<sha>` for source links.
31
+ Use inline SVG or CSS for diagrams, with a text alternative.
32
+
33
+ Count page and reply together under the 700-word cap.
34
+ Use body `textContent` minus script and style, including collapsed content and SVG labels.
35
+ Inline pages have no card exemption.
36
+
37
+ Prefer static pages.
38
+ Interactive means the page has script, form controls, or `<details>`.
39
+ Plain links are not interactive.
40
+ Use semantic elements, labels, visible focus, and sufficient contrast in both themes.
41
+ Pages never use storage, cookies, modals, or popups.
42
+
43
+ ## Check the exact bytes
44
+
45
+ Keep page bytes, SHA-256, and attachmentId in run evidence.
46
+ Scan the exact bytes of every page for remote `src`, `srcset`, and `href` on loaded resources.
47
+ Scan for `@import`, `url()`, `fetch`, XHR, WebSocket, EventSource, and meta refresh.
48
+ Reject scan matches except local `url(#id)` references such as SVG markers and gradients.
49
+
50
+ Before publishing, run `html_preview` at 728px dark and 360px light.
51
+ The driver's two previews are an author exception to delegated browser checks.
52
+ Require correct layout, theme, height, network, word count, zero console errors,
53
+ and empty `missingImages`.
54
+ Check every claim against its source.
55
+ Inspect both screenshots for clipping, overflow, and readable labels.
56
+ For an explicit theme override, also check that requested theme.
57
+
58
+ For every interactive page, require a real-browser `axstack-ui-verifier` pass
59
+ for clicks, keyboard, and expanded states.
60
+ Give the evidence path and SHA-256 of the exact bytes in the verifier brief.
61
+ Follow [UI verification](../../axstack/references/ui-verification.md) for that pass.
62
+ The pass covers focus order, screen-reader names, reduced motion, and an empty
63
+ `performance.getEntriesByType('resource')` list.
64
+
65
+ ## Publish or fall back
66
+
67
+ Publish one page per answer.
68
+ Publish with `html_render` from exactly the checked bytes.
69
+ Fix a failed check and rerun on the new bytes.
70
+ A byte change invalidates affected checks and review.
71
+ If either tool cannot be found after one bounded discovery attempt, fall back
72
+ to chat with Mermaid and state the reason.
73
+ If a check, verifier, or required reviewer is unavailable or still fails, fall
74
+ back to chat with Mermaid and state the reason.
75
+ After an ambiguous `html_render` result, never retry.
76
+ After an ambiguous `html_render` result, read `t3_thread_read`, keep the page
77
+ if present, and otherwise fall back to chat.
@@ -1,5 +1,7 @@
1
1
  # Visual QA checklist
2
2
 
3
+ For inline pages, follow [Inline pages](inline-pages.md) instead of this checklist.
4
+
3
5
  Use this checklist for every HTML explanation and other visual artifacts where
4
6
  rendering matters.
5
7
 
@@ -86,6 +86,11 @@ Size alone never requires user approval.
86
86
 
87
87
  ## 3. Establish test-first evidence
88
88
 
89
+ For performance work only, load [Performance loop](../axstack/references/perf-loop.md).
90
+ For performance work, run the change loop.
91
+ For an accepted optimization scope without new behavior explicitly marked structure-preserving, use the structure-preserving path with the same checks green before and after plus the measured delta.
92
+ For performance work with new behavior, run a failing check first.
93
+
89
94
  Use the normal behavior path unless the accepted improvement scope is
90
95
  explicitly marked **structure-preserving**, or the accepted scope explicitly
91
96
  authorizes **F repairs**. The author never chooses those exceptions.
@@ -167,6 +172,14 @@ evidence through [UI verification](../axstack/references/ui-verification.md)
167
172
  when relevant. Name every unavailable OS, harness, credential, or other
168
173
  boundary instead of implying coverage.
169
174
 
175
+ Use a present `verify-<app>` skill for affected behavior.
176
+ If a PR changes a mapped feature, its author updates the feature file in the same PR.
177
+ For a missing skill, add "suggest `axstack-verify create`" to the receipt's `Unverified:` line.
178
+ Never generate a verification skill automatically.
179
+ For a broken or stale recipe, report the gap.
180
+ Never let a broken or stale recipe waive existing acceptance obligations.
181
+ Keep browser recipes tool-neutral for the verifier.
182
+
170
183
  After the last change, pin the exact candidate revision and return this compact
171
184
  implementation receipt to the driver:
172
185
 
@@ -29,6 +29,10 @@ In mixed fan-out retain a Codex and a Claude seat or hold the affected work.
29
29
 
30
30
  For performance claims only, load [Performance checklist](../axstack/references/performance-checklist.md).
31
31
 
32
+ For performance work only, load [Performance loop](../axstack/references/perf-loop.md).
33
+ For performance work, discovery remains report-only.
34
+ For performance work, rank candidates in mantra order.
35
+
32
36
  ## 1. Bound discovery
33
37
 
34
38
  1. Start with the user's named subsystem or pain. Otherwise inspect recent
@@ -0,0 +1,18 @@
1
+ ---
2
+ name: axstack-perf
3
+ description: When making something faster, use axstack-perf to route measured performance work to the matching phase.
4
+ ---
5
+
6
+ # Performance
7
+
8
+ Usage: `/axstack-perf make the deploy step faster; target -30%`
9
+
10
+ Load [Performance loop](../axstack/references/perf-loop.md).
11
+
12
+ - Route a performance regression to `axstack-debug`.
13
+ - Route optimization discovery to `axstack-improve`.
14
+ - Route an accepted change to `axstack-implement`.
15
+
16
+ Record the target, noise criterion and finite attempt budget as acceptance checks in the routed phase.
17
+ The routed phase owns scope.
18
+ This entry point defines no additional workflow, role or runtime.
@@ -77,6 +77,12 @@ Complete every step before sending.
77
77
  Discovery is complete only when authorization, routing, lookup, and the target
78
78
  listing all pass.
79
79
 
80
+ Before a reply-dependent send, verify the reply route: gateway forwarding to
81
+ `axstack-reply`, the bound inbox path, and the driver's read access.
82
+ If reply-route readiness fails, name this run's T3 driver thread as the action route
83
+ in the message.
84
+ Send it only when the action uses that thread instead of a Telegram reply.
85
+
80
86
  ## Preserve identity and authority
81
87
 
82
88
  Keep every relay body in plain text.
@@ -14,6 +14,11 @@ For every dispatch brief, name its private `<run dir>/evidence/<dispatch>/` fold
14
14
  Produce one user-approved specification whose exact revision can govern
15
15
  ticketing and execution.
16
16
 
17
+ Never add a driver-invented spec decision that makes the user merge PRs outside
18
+ [watch §5](../axstack-watch/SKILL.md#5-state-readiness-precisely)'s user-merge categories.
19
+ Explicit user restrictions, including `Auto-merge: off`, chat holds, and user
20
+ instructions, always win.
21
+
17
22
  Before specification work, load [Standing contracts](../axstack/references/contracts.md).
18
23
  Follow its required path through [Shared lifecycle](../axstack/references/lifecycle.md)
19
24
  and the lifecycle's [audit skill](../axstack-audit/SKILL.md) hook.
@@ -53,18 +58,31 @@ and the lifecycle's [audit skill](../axstack-audit/SKILL.md) hook.
53
58
  substitution. A reviewable draft covers the agreed outcome, acceptance
54
59
  criteria, exclusions, and both adviser receipts or the reported hold.
55
60
  An optional adviser note may be deferred or rejected in a `Decisions` row
56
- with the draft unchanged; it needs no new adviser pair. Changed draft text,
57
- a blocking finding, or a high-stakes decision requires fresh receipts on
58
- the new revision.
61
+ with the draft unchanged; it needs no new adviser pair.
62
+ By default, changed draft text requires fresh receipts on the new revision.
63
+ The only exception to this default is delta confirmation, available only after
64
+ each configured adviser seat confirms the driver's meaning-preservation claim.
65
+ If the driver claims an edit preserves criterion, scope, decision and
66
+ instruction meaning, each configured adviser seat confirms that claim on
67
+ the delta, naming its prior receipt, the delta and the new revision.
68
+ If a seat disagrees, a blocking finding exists, a high-stakes decision arises,
69
+ or meaning changes, obtain full fresh review on the new revision.
70
+ An unavailable adviser seat holds the affected phase.
59
71
  4. **Obtain the specification checkpoint.** The driver owns the draft and the
60
72
  user approves it; adviser input cannot grant approval. High-stakes decisions
61
73
  require `axstack-advisor-astra` and a fresh `axstack-escalation-fable`
62
74
  session to return plain AGREE. Present one
63
75
  reviewable, identified revision for this checkpoint. Its user approval
64
76
  creates the execution baseline.
77
+ In a T3 thread, the driver also publishes the spec readback as an inline page and
78
+ follows [Explain](../axstack-explain/SKILL.md) and
79
+ [Inline pages](../axstack-explain/references/inline-pages.md) for its checks and delivery.
80
+ The readback page shows the revision ID, SHA-256, and store path.
81
+ Approval binds to that revision, not to the page.
65
82
  Name the draft revision covered by each adviser receipt at the checkpoint.
66
83
  If draft text changed and an adviser receipt covers an older revision, hold
67
- approval until fresh receipts cover the presented revision.
84
+ approval until fresh receipts cover the presented revision, using delta
85
+ confirmations only under step 3's meaning-preserving rule.
68
86
  Except for high-stakes decisions, a change confined to a `Decisions` row
69
87
  reuses adviser receipts only while draft text, evidence, scope, and question
70
88
  remain unchanged.
@@ -0,0 +1,151 @@
1
+ ---
2
+ name: axstack-verify
3
+ description: When creating or maintaining a repository verification skill, use axstack-verify to build and prove its recipes through Implement.
4
+ ---
5
+
6
+ # Verification skills
7
+
8
+ Usage: `/axstack-verify create` or `/axstack-verify maintain`
9
+
10
+ Route single-behavior requests to an existing `verify-<app>` skill or
11
+ [UI verification](../axstack/references/ui-verification.md) delegation.
12
+ Never use create for single-behavior requests.
13
+ Deliver create and maintain through `axstack-implement` with one writer and one PR.
14
+ The driver owns scope and dispatch under Implement's standing contracts.
15
+
16
+ ## Create
17
+
18
+ ### Read the repository
19
+
20
+ Read the repository for surface, run, drive, observe and isolate facts:
21
+
22
+ - Surface: identify the primary user interface and record other surfaces.
23
+ - Run: find documented start commands, readiness signals, ports and fixtures.
24
+ - Drive: find stable selectors, commands or endpoints in existing harnesses.
25
+ - Observe: identify action logs, resulting state and side-effect evidence.
26
+ - Isolate: determine how separate instances avoid shared data and port conflicts.
27
+
28
+ Reuse existing harnesses before choosing new tools.
29
+ Ground commands and selectors in this repository.
30
+ For a failing build, report it and hold generation.
31
+ Never edit product code or fix the build.
32
+ If isolation is unavailable, report the gap and hold live driving.
33
+
34
+ ### Write the recipe and map
35
+
36
+ Preserve existing skills.
37
+ Resolve name collisions explicitly before writing.
38
+ Choose a distinct name or return the collision to the driver.
39
+ Write `.agents/skills/verify-<app>/SKILL.md` with Launch, Doctor, Drive, Evidence, Cleanup and Helpers sections.
40
+ Use only `name` and `description` as frontmatter keys.
41
+ Name the app, surface and invocation conditions in the description.
42
+
43
+ - Launch: give exact commands, readiness signals and instance ownership checks.
44
+ - Doctor: give read-only checks for process, version, port ownership and test auth.
45
+ - Drive: give the real user path with stable handles and observable end states.
46
+ Keep browser recipes tool-neutral.
47
+ - Evidence: follow the proof standards in
48
+ [UI verification](../axstack/references/ui-verification.md#proof-standards).
49
+ - Cleanup: stop only processes started by this run and remove owned scratch.
50
+ - Helpers: explain each helper's purpose and working directory.
51
+
52
+ Write `.agents/skills/verify-<app>/features/README.md` as an index of the top 3-5 user features.
53
+ Write per-feature files in `.agents/skills/verify-<app>/features/` and link them from the index.
54
+ Give each feature file its entry points, prerequisites, drive steps and observable proof.
55
+ Record gotchas and uncovered surfaces in the map.
56
+ Commit a relative symlink from `.claude/skills/verify-<app>` to `.agents/skills/verify-<app>`.
57
+ Use `../../.agents/skills/verify-<app>` as its target from `.claude/skills/`.
58
+ Make helpers executable and show each invocation in the skill body.
59
+ Resolve helper paths from the discovered skill directory so the Claude symlink works.
60
+ Test helper logic red/green before implementation.
61
+
62
+ ### Isolate every attempt
63
+
64
+ Launch with an owned 0700 TMPDIR outside HOME for data, home and port bookkeeping.
65
+ Listen on loopback only.
66
+ Never use production secrets.
67
+ Keep evidence in the private evidence folder after cleanup.
68
+ Keep that folder outside scratch selected for deletion.
69
+ Clean up owned resources by literal absolute paths.
70
+ Follow [Safe deletion](../axstack/references/workspace-hygiene.md#safe-deletion).
71
+
72
+ For a peer PR, read the Launch recipe and helpers from the base revision and execute against the pinned candidate.
73
+ Keep trusted base helpers separate from the candidate's recipe files.
74
+ Bind evidence to the served SHA.
75
+ Report an unavailable candidate run as unverified.
76
+
77
+ ### Prove the generated skill
78
+
79
+ Execute launch, doctor, drive one mapped feature, capture evidence, clean up and confirm evidence survives.
80
+ Run doctor before each drive.
81
+ The author drives only non-browser surfaces.
82
+ Delegate browser drives through
83
+ [UI verification](../axstack/references/ui-verification.md) to `axstack-ui-verifier`
84
+ with the feature file path in the brief.
85
+ Give the verifier the pinned candidate, isolated instance and private evidence path.
86
+ Clean up after every failed attempt.
87
+ Retry once after a drift fix.
88
+ If the retry fails, report the failure and hold acceptance.
89
+ Check discovery and relative-helper execution in both Claude Code and Codex.
90
+ If either harness is unavailable, report its checks unverified and hold acceptance.
91
+ Return proof commands, observed outputs, evidence paths and cleanup confirmation.
92
+ Record the bounded exception for generated-skill content on the implement §5 `TDD:` line.
93
+ Keep `axstack-verify` prose-contract red/green tests.
94
+ The exception covers the generated content's executed proof only.
95
+
96
+ ## Maintain
97
+
98
+ Locate the existing `verify-<app>` skill and its feature map.
99
+ Edit only the verification skill's own directory.
100
+ Never edit product code or fix the build.
101
+ Never launch a parallel source wave.
102
+ Never schedule maintain runs.
103
+ Follow [attempt isolation](#isolate-every-attempt) for both modes.
104
+ For a failing build, report it and hold live driving.
105
+
106
+ ### Reconcile the map
107
+
108
+ Check index hygiene against `.agents/skills/verify-<app>/features/README.md` and sibling files for missing, extra, duplicate and dead entries.
109
+ Read each mapped feature from source and record entry points with citations.
110
+ Report omitted entry points with source paths separately from passes.
111
+ Classify each gap as doc drift, harness gap or product regression:
112
+
113
+ - Doc drift: the map differs from intended source behavior.
114
+ For doc drift, fix the map.
115
+ - Harness gap: working behavior cannot be driven by the recipe.
116
+ For a harness gap, fix the recipe.
117
+ - Product regression: the app fails its intended behavior.
118
+ For a product regression, report it to the driver, which routes it to `axstack-debug`.
119
+ Never fix a product regression in docs.
120
+
121
+ ### Prove the map
122
+
123
+ Drive every mapped feature live even when source looks clean.
124
+ Use the skill's Launch recipe for an isolated instance.
125
+ Run doctor before each drive.
126
+ The author drives only non-browser surfaces.
127
+ Delegate browser drives through
128
+ [UI verification](../axstack/references/ui-verification.md) to `axstack-ui-verifier`
129
+ with the feature file path in the brief.
130
+ Capture actions, resulting states and side effects under its proof standards.
131
+ Report unreachable prerequisites with attempted routes separately from passes.
132
+ Fix a missing prerequisite in the map as doc drift.
133
+ Clean up after every failed attempt.
134
+ Retry once after a drift fix.
135
+ If the retry fails, report `blocked` and hold acceptance.
136
+ Re-drive each corrected recipe live before returning it.
137
+ Keep helpers executable with their invocation documented in the skill body.
138
+ After the final drive, clean up owned resources.
139
+ Confirm evidence survives each cleanup in the private evidence folder.
140
+
141
+ ### Return the outcome
142
+
143
+ Report `clean` only when every mapped feature has source and live coverage with successful drives and zero corrections.
144
+ Report `changed` for one PR of proven corrections.
145
+ Report `blocked` when coverage cannot finish or corrections cannot ship safely.
146
+ Never report incomplete mapped coverage as `clean`.
147
+ Return the feature coverage, gaps, proof commands, evidence paths and cleanup confirmation in the implementation receipt.
148
+
149
+ Ideas paraphrased from pstack's
150
+ [create-verification-skill](https://github.com/cursor/plugins/blob/d0ef80d86795816da932a153458c5dbe192d294e/pstack/skills/create-verification-skill/SKILL.md) (MIT) and
151
+ [maintain-verification-skill](https://github.com/cursor/plugins/blob/d0ef80d86795816da932a153458c5dbe192d294e/pstack/skills/maintain-verification-skill/SKILL.md) (MIT).
@@ -251,6 +251,7 @@ Never auto-merge PRs changing lockfiles.
251
251
  Never auto-merge PRs changing test-runner config.
252
252
  Never auto-merge PRs changing branch-protection or ruleset config.
253
253
  Never auto-merge PRs changing `CODEOWNERS`.
254
+ Never auto-merge PRs changing `AGENTS.md`.
254
255
  Test sources stay eligible.
255
256
  Axstack skill and merge-rule text are eligible under the watch predicate.
256
257
  Changes to package.json limited to `version` and `files` are eligible under
@@ -280,7 +281,7 @@ In `team` mode the reply only clears an ineligible base, auto-merge turned off,
280
281
  and an open human or bot comment.
281
282
  PRs in the CI, package.json beyond `version` and `files`, lockfile, test-runner,
282
283
  branch-protection and ruleset,
283
- `CODEOWNERS`, or non-`clean` revert categories are merged by the user on the forge.
284
+ `CODEOWNERS`, `AGENTS.md`, or non-`clean` revert categories are merged by the user on the forge.
284
285
  Promotion, `deploying`-base, unknown-base, and peer PRs are merged by the user on the
285
286
  forge, and the card only reports readiness.
286
287
  User merges are bottom-up for a stack.
@@ -42,8 +42,11 @@ One bound schedule serves both the run watch and the chat-run watch; never creat
42
42
  The chat-run watch never expires or waits for re-authorization while PRs remain open.
43
43
  If the native schedule has a lifetime, the driver re-arms it at a wake.
44
44
  Use `update_scheduled_task` on the recorded schedule ID for cadence changes and re-arming.
45
- After 7 days with no event on any watched PR, and only with no unsettled launched work,
46
- change the wake cadence from 5 to 60 minutes.
45
+ When the run waits only on a user decision or a human-only step, including a
46
+ user merge, with no unsettled worker and no PR needing watch events, change
47
+ the run watch cadence to 60 minutes at once.
48
+ When work restarts, restore the normal 5-minute cadence.
49
+ Keep native PR watches.
47
50
  On the next event on a watched PR, restore the wake cadence to 5 minutes.
48
51
  If launched work becomes unsettled, restore the 5-minute cadence.
49
52
  Read back each schedule update and record its receipt; an uncertain update holds affected work.