axstack 0.22.0 → 0.23.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -23,6 +23,7 @@ coordination. You can start at the phase you need.
23
23
  | Verify | [axstack-review](skills/axstack-review/SKILL.md) | Review a PR or bounded codebase at an exact revision. |
24
24
  | Verify | [axstack-improve](skills/axstack-improve/SKILL.md) | Find evidenced codebase improvements without editing code. |
25
25
  | Verify | [axstack-audit](skills/axstack-audit/SKILL.md) | Measure a run's outcomes and evidence gaps. |
26
+ | Verify | [axstack-correct](skills/axstack-correct/SKILL.md) | Report repeated mistakes and propose stronger checks when invoked by the user. |
26
27
  | Operate | [axstack-watch](skills/axstack-watch/SKILL.md) | Observe or maintain an existing PR within its authority. |
27
28
  | Operate | [axstack-cleanup](skills/axstack-cleanup/SKILL.md) | Retire eligible completed agent resources. |
28
29
  | Operate | [axstack-relay](skills/axstack-relay/SKILL.md) | Send an explicit message or authorized notification. |
package/docs/workflows.md CHANGED
@@ -18,6 +18,8 @@ and PR shape.
18
18
  Direct routes need no spec ceremony:
19
19
 
20
20
  - `axstack-research` answers one bounded source-backed question.
21
+ - `axstack-correct` reports repeated mistakes and proposes stronger checks.
22
+ Only the user invokes it.
21
23
  - `axstack-explain` separates implemented, intended, tested, live, and unknown
22
24
  behavior; complex visuals receive exact-artifact QA where applicable.
23
25
  - `axstack-improve` returns a small ranked set of evidenced improvement
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "axstack",
3
- "version": "0.22.0",
3
+ "version": "0.23.0",
4
4
  "description": "Axstack installer and setup CLI: installs owned chat skills and role data, configures supported harness settings, and checks T3 Code capabilities.",
5
5
  "keywords": [
6
6
  "claude-code",
@@ -17,7 +17,16 @@ protection; a mismatch holds publication. If the push outcome is ambiguous,
17
17
  inspect remote state before retrying.
18
18
 
19
19
  Before reviewer dispatch, read the remote ref back and confirm that it resolves
20
- to the candidate SHA; also pin the current base. Record:
20
+ to the candidate SHA; also pin the current base.
21
+ After every push, before post-push diligence or merge-ready, compare the PR body's
22
+ stated head SHA with `Confirmed remote SHA` and record `PR body head SHA: <sha or none>`.
23
+ Accept `none` for a PR body without a stated head SHA.
24
+ Accept a stated full head SHA only when it equals `Confirmed remote SHA`.
25
+ Accept a stated abbreviated head SHA only when it is a matching prefix of
26
+ `Confirmed remote SHA` with at least 7 hexadecimal characters.
27
+ Hold on a mismatched head SHA or an abbreviation shorter than 7 characters.
28
+ Leave other SHAs in the PR body unchanged.
29
+ Record:
21
30
 
22
31
  ```text
23
32
  Candidate: <sha>
@@ -25,6 +34,7 @@ Base: <sha>
25
34
  Remote ref: <branch>
26
35
  Expected-old remote SHA: <sha | absent>
27
36
  Confirmed remote SHA: <sha>
37
+ PR body head SHA: <sha or none>
28
38
  PR: <url>
29
39
  CI: <run ID or URL and triggered/pending/completed status>
30
40
  ```
@@ -2,10 +2,12 @@
2
2
 
3
3
  Apply these authority, scope, and model rules before consequential action.
4
4
 
5
+ For user-facing output, follow [STE-inspired writing](ste-writing.md).
6
+
5
7
  ## Required lifecycle load
6
8
 
7
- Except for `axstack-audit` and `axstack-relay`, every independently called phase
8
- must load and follow [Shared lifecycle](lifecycle.md) before acting. When a
9
+ Except for `axstack-audit`, `axstack-relay`, and `axstack-correct`, every
10
+ independently called phase must load and follow [Shared lifecycle](lifecycle.md) before acting. When a
9
11
  substantive run ends or reaches a meaningful checkpoint, apply the lifecycle
10
12
  audit hook. The audit phase loads these contracts, writes its assigned record,
11
13
  and stops; it never audits itself.
@@ -3,7 +3,7 @@
3
3
  Read the [T3 runtime boundary](t3-runtime.md) before dispatch.
4
4
  Dispatch `axstack-diligence` with async `delegate_task` in the driver worktree,
5
5
  using a pinned brief and private evidence paths.
6
- It is read-only, never authors or edits, and returns `PASS` or `FINDINGS`
6
+ It is read-only, never authors or edits, and returns `PASS`, `FINDINGS`, or `UNKNOWN`
7
7
  with locations, observed evidence, and limits. A stale or missing receipt is
8
8
  not a pass. Keep its first pass independent of other reviewers and workers.
9
9
 
@@ -21,5 +21,14 @@ publication, compare the author receipt with its evidence folder: red/green
21
21
  logs exist, and counts, SHAs, and paths match. For release preparation, compare
22
22
  the release PR body with the merged PRs.
23
23
 
24
+ Retain the full output of every full-suite run as a named log in the dispatch's
25
+ private evidence folder.
26
+ For an unattributed full-suite failure, report `UNKNOWN` and return the failure to the driver.
27
+ Never report `PASS` for an unattributed full-suite failure.
28
+ The driver records the disposition of each full-suite failure in the run record.
29
+ Report an attributed failure with its retained output normally.
30
+ An observed full-suite failure still fails the suite, including when attributed or reported `UNKNOWN`.
31
+ No reruns are required.
32
+
24
33
  `FINDINGS` identifies a mismatch for the driver to resolve at the owning phase;
25
34
  it does not edit the artifact or create another review round by itself.
@@ -118,6 +118,6 @@ After preservation and salvage checks, archive the exact eligible T3 thread
118
118
  with `t3_thread_organize`, then remove its exact recorded checkout path using
119
119
  `git worktree remove <path>` without force. Verify absence with
120
120
  `git worktree list --porcelain`; retire only eligible local-only branches with
121
- `git branch -d` under workspace hygiene. Never use shell recursive deletion or
121
+ `git branch -d` under workspace hygiene. Never use shell recursive deletion to remove a whole worktree or
122
122
  treat archive success as ownership, settlement, liveness or cleanup proof.
123
123
  Failure or uncertainty preserves the resource.
@@ -143,6 +143,8 @@ nothing without tested independent review.
143
143
 
144
144
  ## Close-out
145
145
 
146
+ Close-out requires the [close-out acceptance table](run-record.md#close-out-acceptance).
147
+
146
148
  PRs merge by forge state; close out: (1) settle every T3 worker run and archive eligible threads; (2) compact record with counts and denominators—user
147
149
  interventions/deviations from plan/repairs; (3) `axstack-auditor`: settle
148
150
  non-zero/requested, else `counts zero`. A base auditor preflight rejection
@@ -153,7 +155,7 @@ use `axstack-cleanup`, remove the run's own scheduled tasks under
153
155
  [Workspace hygiene](workspace-hygiene.md), and close selected external-tracker tickets;
154
156
  (5) mark the
155
157
  [Run record](run-record.md) `Archived`. `Archived`—one each:
156
- settlement receipt; compact record path; auditor decision plus settlement
158
+ settlement receipt; compact record path; close-out acceptance table; auditor decision plus settlement
157
159
  receipt, `counts zero`, or the unlaunchable UNKNOWN archive receipt; scheduled-task,
158
160
  release, and ticket receipts; archive timestamp.
159
161
  `active`/receipt-incomplete record: close-out pending, never done. One-step
@@ -0,0 +1,22 @@
1
+ # Performance checklist
2
+
3
+ Use this checklist only for performance claims.
4
+ Keep workflow count accounting with axstack-audit.
5
+ A limiter is the resource or code path that bounds performance.
6
+
7
+ 1. **Limiter: What bounds the result?**
8
+ Identify the resource or code path that limits the result from profiles or counters captured during a run.
9
+ 2. **Tuning: Did each option use suitable settings?**
10
+ Check each option with comparable production settings, versions, and data.
11
+ 3. **Limits: Is the result physically plausible?**
12
+ Compare the result with hardware limits and the changed code's share of total elapsed time.
13
+ 4. **Errors: Did the run succeed?**
14
+ Verify successful work and correct outputs using error counts from the measurement run.
15
+ 5. **Reproducibility: Does the difference repeat?**
16
+ Report the median and range from repeated, alternating runs of each option.
17
+ 6. **Relevance: Does the result matter to users?**
18
+ Measure the end-to-end user path with realistic data and concurrency.
19
+ 7. **Work happened: Did the timed work finish?**
20
+ Verify that the intended work completed inside the timed region.
21
+
22
+ Ideas paraphrased from [pstack benchmark-checklist](https://github.com/cursor/plugins/blob/e43c7ee26e0038c6c1fa8380dd34ce86ff94cb2a/pstack/skills/benchmark-checklist/SKILL.md) (MIT), using Brendan Gregg's seven benchmark questions.
@@ -65,6 +65,8 @@ step (3) for user routing: no substitution or same-provider review.
65
65
  off a classified repair (explain: how; debug: what's wrong).
66
66
  - Code quality/refactor discovery -> `axstack-improve`: rank bounded
67
67
  candidates with evidence; report only, no source edits.
68
+ - Repeated mistakes need evidence and stronger checks -> `axstack-correct`:
69
+ user-invoked, report only.
68
70
  - Accepted worker/task/run completion or bounded backlog request -> driver invokes
69
71
  `axstack-cleanup` inline; never dispatch it.
70
72
  - Preparation completion, watch expiry, resume, or reconciliation -> the
@@ -119,6 +119,21 @@ by the next owner:
119
119
  evidence lives. Read it on resume before reconciling; append, never rewrite,
120
120
  and keep entries as short as the evidence pointer allows.
121
121
 
122
+ ## Close-out acceptance
123
+
124
+ Use one row per acceptance check, including each clause of a compound check.
125
+ Record in each row a passing evidence pointer or a user-accepted hold with its `Decisions` row.
126
+ If neither passing evidence nor a `Decisions` row with a user-accepted hold exists for an acceptance clause, hold close-out.
127
+ A recorded user-accepted hold in `Decisions` satisfies that clause for close-out; keep the unmet result explicit.
128
+ Record the reason for each driver-elected repair in `Decisions`.
129
+
130
+ ```markdown
131
+ | Acceptance check | Result | Evidence or user-accepted hold | Repair election (Decisions row + reason, or none) |
132
+ | --- | --- | --- | --- |
133
+ | <check + clause> | passed | <SHA + check/log pointer> | <Decisions row + reason, or none> |
134
+ | <check + unmet clause> | held (user accepted) | <Decisions row + user acceptance receipt> | <Decisions row + reason, or none> |
135
+ ```
136
+
122
137
  ## Privacy
123
138
 
124
139
  Record concise IDs, SHAs, URLs, status, timestamps, next actions, and evidence
@@ -0,0 +1,22 @@
1
+ # STE-inspired writing
2
+
3
+ STE means Simplified Technical English.
4
+ This reference borrows clarity rules from STE.
5
+
6
+ Follow these rules only for new or materially revised user-facing output.
7
+ User-facing output includes reports, PR bodies, briefs, read-backs, and packaged guidance.
8
+ Do not rewrite existing prose for style.
9
+ Keep exact identifiers unchanged.
10
+ Keep quotes unchanged.
11
+ Keep safety contracts unchanged.
12
+ Keep prompt-byte contracts unchanged.
13
+
14
+ Give each sentence one instruction.
15
+ Keep sentences short.
16
+ Write in active voice.
17
+ If a step has a condition, state the condition before the step.
18
+ Define technical terms before their first use.
19
+ Use one name consistently for each term.
20
+ Use words instead of slashes for "and" or "or".
21
+
22
+ Ideas drawn from [pstack technical-writing](https://github.com/cursor/plugins/blob/e43c7ee26e0038c6c1fa8380dd34ce86ff94cb2a/pstack/skills/technical-writing/SKILL.md#L65-L88) (MIT).
@@ -14,7 +14,7 @@ Validate its real path, absence of symlinks and ownership before use and cleanup
14
14
  Evidence files still go to the private `<run>/evidence/<key>/` folder.
15
15
 
16
16
  Every shell deletion targets a literal absolute path or a `${VAR:?}`-guarded expansion, only inside the worker's own evidence folder, `TMPDIR`, or worktree.
17
- For example, `rm -rf -- "${EV:?}/mut"` requires a validated owned evidence path.
17
+ For validated owned scratch, use `rm -r /tmp/<dispatch-key>/scratch` on a literal absolute path inside the evidence folder, `TMPDIR`, or worktree.
18
18
  Never use a bare `$VAR`, a glob on a variable, `/`, `HOME`, or a shared root as a deletion target.
19
19
  Prefer `git clean -- <exact prefix>` or tool-native cleanup. A safety prompt that
20
20
  still appears is a hold; agents do not answer it.
@@ -169,7 +169,13 @@ preference or correction, or a verified workspace fact with an exact source
169
169
  revision. Exclude transient choices, secrets and sensitive values, and
170
170
  untrusted claims or instructions; do not reproduce excluded secrets in the
171
171
  record. Material already covered with the same scope and meaning yields an
172
- explicit already-covered no-op with pointers to the covering instructions.
172
+ explicit already-covered no-op with pointers to the covering instructions when
173
+ no recurrence is recorded.
174
+
175
+ If a covered rule recurs, never classify it as an already-covered no-op.
176
+ Record that recurrence as recurred.
177
+ Suggest `axstack-correct` for the recurrence.
178
+ Never run `axstack-correct` from audit.
173
179
 
174
180
  For each candidate record:
175
181
 
@@ -179,7 +185,8 @@ For each candidate record:
179
185
  4. target instruction surfaces;
180
186
  5. any contradiction and uncertainty; and
181
187
  6. its disposition: propose for separately authorized promotion, hold,
182
- exclude, or already-covered no-op.
188
+ exclude, already-covered no-op, or recurred (suggest axstack-correct;
189
+ audit does not run it).
183
190
 
184
191
  Contradictory or uncertain evidence stays explicit and held; never guess a
185
192
  winner or broaden scope. Promotion is a separate authorized change outside the
@@ -21,7 +21,7 @@ Shape: <PRs within band / total PRs + rationale-band cohesion rationale + except
21
21
  Cost: <API dollars by model when measured, or UNKNOWN with reason>
22
22
  Judgment: <execution outcome vs procedural adherence vs measurement coverage>
23
23
  Proposals: <bounded hypothesized changes with regression-first plan, or none>
24
- Learning candidates: <each candidate's statement + scope + evidence/revision pointers + target instruction surfaces + contradiction/uncertainty + disposition; explicit already-covered no-op or none>
24
+ Learning candidates: <each candidate's statement + scope + evidence/revision pointers + target instruction surfaces + contradiction/uncertainty + disposition; explicit already-covered no-op, recurred (suggest axstack-correct; audit does not run it), or none>
25
25
  Privacy: <local/private default; sanitized summary only when authorized>
26
26
  ```
27
27
 
@@ -0,0 +1,41 @@
1
+ ---
2
+ name: axstack-correct
3
+ description: When repeated mistakes need evidence and stronger checks, use axstack-correct to report correction proposals.
4
+ disable-model-invocation: true
5
+ ---
6
+
7
+ # Correct
8
+
9
+ Usage: /axstack-correct ["<correction>"] [runs=<ids> | last=<N, default 10>]
10
+
11
+ Load [Standing contracts](../axstack/references/contracts.md) before acting.
12
+
13
+ Run only at the user's request.
14
+ Write a report only.
15
+ Never edit files.
16
+ Never dispatch workers.
17
+ Never merge.
18
+ Never read transcripts.
19
+
20
+ Read [run records](../axstack/references/run-record.md): Decisions, deviations, Learnings, FAILED receipts, audit proposals.
21
+ Read git history bounded by selected runs.
22
+ Read only PR reviews named in those records.
23
+ Record selected runs and git bounds.
24
+ Report inaccessible evidence.
25
+
26
+ A class needs at least two distinct incidents.
27
+ Each incident needs a distinct pointer: run id + attempt or file:line, SHA, or PR or review URL.
28
+ Count an incident echoed in several records once.
29
+ If a pointer is missing, report recurrence UNKNOWN and keep the class open.
30
+
31
+ Try fixes in this order: remove copied pattern > script or helper error naming the fix > brief or receipt template > bun test > prose.
32
+ Label each rule `shipped` or `runtime: advisory`.
33
+ Use `shipped` only when a check fails in CI on the mistake itself.
34
+ When only audit can observe a rule, label it `runtime: advisory`.
35
+ A prose-contract test alone leaves a rule `runtime: advisory`.
36
+ Propose fixes through a small-change intent and [axstack-implement](../axstack-implement/SKILL.md).
37
+
38
+ The `## Enforced rules` table in the target repo's AGENTS.md changes only in the PR that adds its check.
39
+ The first row's PR adds a test that fails when any row's enforcement path disappears.
40
+ If a `shipped` check missed a later recorded recurrence, report `enforcement failed`.
41
+ For that failure, propose relabelling the row `runtime: advisory` in a separately reviewed PR.
@@ -22,6 +22,8 @@ read it before phase 7 and before any fan-out.
22
22
  Redact secrets before showing any command, output, or artifact; build loops
23
23
  against environment variables so credentials never appear in what is shown.
24
24
 
25
+ For performance claims only, load [Performance checklist](../axstack/references/performance-checklist.md).
26
+
25
27
  ## Phases
26
28
 
27
29
  Each phase has an observable completion criterion. Skip one only with a
@@ -230,6 +230,11 @@ For each PR:
230
230
  to step 1.
231
231
  Merge-ready also requires a current diligence `PASS` at that head; diligence
232
232
  `FINDINGS` return to the same author within the review round.
233
+ Diligence `UNKNOWN` records the PR as `held` with the reason in the run record
234
+ and blocks merge-ready pending the driver's recorded disposition.
235
+ Before merge-ready, treat a run-record `Learning` that contradicts a shipped
236
+ rule as a finding on the owning PR and hold merge-ready until the contradiction is resolved.
237
+ An unrelated run-record `Learning` leaves merge-ready eligibility unchanged.
233
238
  A round with reviewer `REQUEST_CHANGES` and/or diligence `FINDINGS` increments
234
239
  `repairs` once and counts once toward the third-round hold.
235
240
 
@@ -27,6 +27,8 @@ Sol pair that fails to launch is fenced, recorded `absent (<reason>)`, and
27
27
  named once in the next read-back, then skipped without relay or substitution.
28
28
  In mixed fan-out retain a Codex and a Claude seat or hold the affected work.
29
29
 
30
+ For performance claims only, load [Performance checklist](../axstack/references/performance-checklist.md).
31
+
30
32
  ## 1. Bound discovery
31
33
 
32
34
  1. Start with the user's named subsystem or pain. Otherwise inspect recent
@@ -33,6 +33,8 @@ When the caller is a bounded review-manager PR job, load
33
33
  carry the required escalation field and every eligible peer PR takes a binding
34
34
  `APPROVE` or `REQUEST_CHANGES` verdict under the automation exception below.
35
35
 
36
+ For performance claims only, load [Performance checklist](../axstack/references/performance-checklist.md).
37
+
36
38
  ## Codebase findings mode
37
39
 
38
40
  Use this manual mode for existing code at a pinned exact source revision and a
@@ -62,6 +62,12 @@ and the lifecycle's [audit skill](../axstack-audit/SKILL.md) hook.
62
62
  session to return plain AGREE. Present one
63
63
  reviewable, identified revision for this checkpoint. Its user approval
64
64
  creates the execution baseline.
65
+ Name the draft revision covered by each adviser receipt at the checkpoint.
66
+ If draft text changed and an adviser receipt covers an older revision, hold
67
+ approval until fresh receipts cover the presented revision.
68
+ Except for high-stakes decisions, a change confined to a `Decisions` row
69
+ reuses adviser receipts only while draft text, evidence, scope, and question
70
+ remain unchanged.
65
71
  Before user approval, dispatch `axstack-diligence` under
66
72
  [Diligence](../axstack/references/diligence.md) to check the draft against
67
73
  the Align decisions for anything dropped, added, or softened.