axstack 0.20.31 → 0.21.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +22 -21
- package/bin/axstack.js +17 -5
- package/docs/installation.md +97 -48
- package/docs/workflows.md +165 -122
- package/package.json +3 -3
- package/profiles/presets/claude-only.json +23 -23
- package/profiles/presets/codex-only.json +10 -10
- package/profiles/presets/mixed.json +24 -24
- package/skills/axstack/references/automations.md +127 -137
- package/skills/axstack/references/autopilot.md +30 -17
- package/skills/axstack/references/candidate-publication.md +13 -8
- package/skills/axstack/references/contracts.md +10 -8
- package/skills/axstack/references/diligence.md +3 -1
- package/skills/axstack/references/evidence-archive.md +38 -33
- package/skills/axstack/references/lifecycle.md +64 -50
- package/skills/axstack/references/review-manager-prompt.md +13 -11
- package/skills/axstack/references/role-roster.md +12 -2
- package/skills/axstack/references/routing.md +29 -25
- package/skills/axstack/references/run-record.md +35 -16
- package/skills/axstack/references/t3-runtime.md +234 -0
- package/skills/axstack/references/test-audit-weekly.md +62 -0
- package/skills/axstack/references/test-value.md +120 -0
- package/skills/axstack/references/ui-verification.md +5 -1
- package/skills/axstack/references/workspace-hygiene.md +102 -156
- package/skills/axstack/scripts/pr-digest.js +120 -0
- package/skills/axstack/scripts/resolve-models.js +94 -38
- package/skills/axstack-align/SKILL.md +17 -7
- package/skills/axstack-audit/SKILL.md +12 -3
- package/skills/axstack-audit/references/record.md +1 -1
- package/skills/axstack-cleanup/SKILL.md +69 -87
- package/skills/axstack-debug/SKILL.md +1 -1
- package/skills/axstack-explain/SKILL.md +1 -1
- package/skills/axstack-explain/references/visual-qa.md +2 -0
- package/skills/axstack-implement/SKILL.md +56 -20
- package/skills/axstack-improve/SKILL.md +24 -4
- package/skills/axstack-relay/SKILL.md +8 -6
- package/skills/axstack-research/SKILL.md +11 -4
- package/skills/axstack-review/SKILL.md +34 -30
- package/skills/axstack-spec/SKILL.md +18 -13
- package/skills/axstack-tickets/SKILL.md +7 -8
- package/skills/axstack-watch/SKILL.md +97 -27
- package/skills/axstack-watch/references/watch-runtime.md +51 -66
- package/src/capabilities.js +33 -69
- package/src/installer.js +1 -1
- package/src/instructions.js +9 -4
- package/skills/axstack/references/orca-runtime.md +0 -202
- package/skills/axstack/scripts/trust-path.js +0 -123
|
@@ -23,7 +23,7 @@ shared load edge explicit: Standing contracts require
|
|
|
23
23
|
substantive phases, and lifecycle's audit hook loads this skill. This audit is
|
|
24
24
|
the terminal exception: it writes its assigned record and does not audit itself.
|
|
25
25
|
|
|
26
|
-
The dispatching driver
|
|
26
|
+
The dispatching driver must read [T3 runtime](../axstack/references/t3-runtime.md)
|
|
27
27
|
immediately before an actual auditor profile or session dispatch. Ordinary
|
|
28
28
|
audit reading and record writing do not load it, and the auditor never
|
|
29
29
|
dispatches.
|
|
@@ -34,7 +34,14 @@ This skill governs what that auditor reads, measures, and proposes.
|
|
|
34
34
|
Dispatch `axstack-auditor` and `axstack-auditor-sol` independently on the same
|
|
35
35
|
bounded brief, without cross-reading. The driver reconciles findings per claim;
|
|
36
36
|
never average verdicts. Record an intentionally absent Sol seat and continue
|
|
37
|
-
with the base auditor alone
|
|
37
|
+
with the base auditor alone. `axstack-auditor-sol` is optional: if its launch
|
|
38
|
+
fails, fence it, record `absent (<reason>)`, name it once in the next read-back,
|
|
39
|
+
and skip it without relay or substitution. In mixed fan-out retain a Codex and
|
|
40
|
+
a Claude seat or hold the affected audit.
|
|
41
|
+
If the base auditor is unlaunchable (preflight rejection, no Dispatch started),
|
|
42
|
+
record `auditor: UNKNOWN (unlaunchable)` with the attempted route and error as
|
|
43
|
+
the archive receipt; archive the run. A launched auditor Dispatch must settle
|
|
44
|
+
normally. There is no substitution for the base auditor.
|
|
38
45
|
The user-chosen improvement mode is a tested, independently reviewed PR that a
|
|
39
46
|
human merges.
|
|
40
47
|
|
|
@@ -87,7 +94,9 @@ counts with denominators plus the evidence behind the count:
|
|
|
87
94
|
- Applicable test-first evidence: normal behavior changes have real red-green
|
|
88
95
|
proof; explicitly accepted structure-preserving work has the old revision
|
|
89
96
|
green before edits and the same checks green on the new revision, plus
|
|
90
|
-
applicable equivalence evidence.
|
|
97
|
+
applicable equivalence evidence. Authorized F repairs use
|
|
98
|
+
[F proof](../axstack/references/test-value.md#f-proof).
|
|
99
|
+
Record noncompliance when the applicable
|
|
91
100
|
evidence path is absent, or `UNKNOWN` with the reason when its records are
|
|
92
101
|
unavailable.
|
|
93
102
|
- Independent exact-revision review status and unresolved findings.
|
|
@@ -12,7 +12,7 @@ Steps: <completed / deviated + why + approval per deviation>
|
|
|
12
12
|
Advisers: <Astra/Opus configured coverage / eligible uses + same-question receipts; Astra/Fable high-stakes AGREE coverage / eligible uses>
|
|
13
13
|
Decisions: <escalation trigger + evidence pointers + outcome changed yes/no, or n/a>
|
|
14
14
|
Debug: <rung reached + loop command + fix attempts + adviser and investigator receipts + isolation evidence | n/a>
|
|
15
|
-
TDD: <applicable evidence path: normal real red-green | accepted structure-preserving old revision green before edits + same checks new revision green; absent proof: noncompliance | unavailable records: UNKNOWN with reason>
|
|
15
|
+
TDD: <applicable evidence path: normal real red-green | accepted structure-preserving old revision green before edits + same checks new revision green | F-repair base-green/removal-inversion-red/rewording-green evidence; absent proof: noncompliance | unavailable records: UNKNOWN with reason>
|
|
16
16
|
Review: <exact-rev independent review status + unresolved findings>
|
|
17
17
|
Rework: <cycles + causes>
|
|
18
18
|
Interventions: <avoidable user interventions, or unsupported by records>
|
|
@@ -1,11 +1,11 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: axstack-cleanup
|
|
3
|
-
description: When completed
|
|
3
|
+
description: When completed T3 threads and worktrees need bounded retirement, use axstack-cleanup after accepted settlement or for an explicitly scoped backlog.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Cleanup
|
|
7
7
|
|
|
8
|
-
Run cleanup inline in the driver after accepting a worker
|
|
8
|
+
Run cleanup inline in the driver after accepting a worker task or run
|
|
9
9
|
completion, or for the exact backlog scope the user named. This skill never
|
|
10
10
|
dispatches a cleanup worker and never retires its current driver session.
|
|
11
11
|
|
|
@@ -14,13 +14,15 @@ Before any runtime action, load and follow:
|
|
|
14
14
|
- [Standing contracts](../axstack/references/contracts.md)
|
|
15
15
|
- [Lifecycle and receipts](../axstack/references/lifecycle.md)
|
|
16
16
|
- [Shared routing](../axstack/references/routing.md)
|
|
17
|
-
- [
|
|
17
|
+
- [T3 runtime](../axstack/references/t3-runtime.md)
|
|
18
18
|
- [Workspace hygiene](../axstack/references/workspace-hygiene.md) for settlement, salvage, and driver-start sweep
|
|
19
19
|
- [Private evidence archive](../axstack/references/evidence-archive.md) when
|
|
20
20
|
evidence is the last removable-worktree blocker
|
|
21
21
|
|
|
22
|
-
Use
|
|
23
|
-
|
|
22
|
+
Use project-scoped T3 thread inventory, `git worktree list --porcelain` and run
|
|
23
|
+
records. Preflight requires `worktreeCleanup` off via `t3_project_read` where
|
|
24
|
+
exposed, or a recorded setup limitation. Follow the runtime schema, never invent
|
|
25
|
+
an alternate command protocol.
|
|
24
26
|
|
|
25
27
|
## Authority and scope
|
|
26
28
|
|
|
@@ -28,29 +30,26 @@ Inline cleanup may consider only resources owned by the accepted completion it
|
|
|
28
30
|
is processing. Driver-start orphan sweeps follow the guarded cross-run sweep in
|
|
29
31
|
[Workspace hygiene](../axstack/references/workspace-hygiene.md). Backlog
|
|
30
32
|
cleanup requires an explicit bounded selector such as a
|
|
31
|
-
|
|
33
|
+
project, run, task set, checkout set, repository, or named age window; age narrows an
|
|
32
34
|
inventory but never establishes eligibility. A partial inventory holds only the
|
|
33
35
|
resource whose identity or state is incomplete while other independently proven
|
|
34
36
|
resources may proceed.
|
|
35
37
|
|
|
36
|
-
Never clean a
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
evidence that has not been durably preserved. Do not
|
|
41
|
-
force native removal, bulk-clean, override a hook failure, edit a runtime
|
|
42
|
-
database, or add a scheduler, daemon, or state machine.
|
|
38
|
+
Never clean a user-created thread, the current driver, a user-taken-over thread, an active or unknown worker, an unsettled descendant, or a resource with ambiguous ownership.
|
|
39
|
+
Preserve unknown files, unmerged author work, ambiguous publication, and evidence that has not been durably preserved.
|
|
40
|
+
Do not force removal, bulk-clean, edit a runtime database, or add a scheduler,
|
|
41
|
+
daemon or state machine. Report other projects' resources; never touch them.
|
|
43
42
|
|
|
44
43
|
## Reconcile each candidate
|
|
45
44
|
|
|
46
|
-
Take a fresh native inventory and bind every candidate to its exact
|
|
47
|
-
|
|
45
|
+
Take a fresh native inventory and bind every candidate to its exact projectId, attempt key, taskId/childThreadId/childRunId
|
|
46
|
+
or threadId/runId, checkout path, repository, branch, and current
|
|
48
47
|
liveness. Read current Git and forge state rather than trusting age, names, or a
|
|
49
48
|
prior receipt. Reconcile an existing cleanup claim before retrying so repeated
|
|
50
49
|
invocations converge instead of duplicating mutations.
|
|
51
50
|
|
|
52
51
|
Retire descendants before parents. A candidate is eligible only when all owned
|
|
53
|
-
|
|
52
|
+
tasks and runs are accepted as settled, no descendant remains unsettled, native
|
|
54
53
|
liveness is positively known where required, and every preservation guard is
|
|
55
54
|
cleared. Record one decision per resource; uncertainty about one candidate does
|
|
56
55
|
not authorize or block unrelated candidates.
|
|
@@ -60,50 +59,47 @@ not authorize or block unrelated candidates.
|
|
|
60
59
|
Classify exact evidence files individually. Save the compact cleanup decision
|
|
61
60
|
and identities in the private run record or another configured durable private
|
|
62
61
|
location outside disposable worktrees. When the evidence archive applies, use
|
|
63
|
-
its helper with either the existing PR identity or the non-PR
|
|
62
|
+
its helper with either the existing PR identity or the non-PR run and task
|
|
64
63
|
identity; never invent a PR number. Read back the durable record and, when an
|
|
65
64
|
archive is used, its manifest, including hashes and exact identities, before
|
|
66
65
|
removing any source copy or workspace.
|
|
67
66
|
|
|
68
|
-
Archive success proves only preservation of the listed bytes
|
|
69
|
-
settlement, exit, ownership, a clean worktree, publication, or removal safety.
|
|
67
|
+
Archive success proves only preservation of the listed bytes; it does not prove settlement, liveness, ownership, a clean worktree, publication, or removal safety.
|
|
70
68
|
|
|
71
69
|
For a completed non-author worktree with useful local content, follow the
|
|
72
70
|
[Workspace hygiene](../axstack/references/workspace-hygiene.md) salvage path
|
|
73
71
|
before removal; a verified bundle changes preservation classification, not
|
|
74
72
|
native ownership or liveness. Keep an author worktree until its PR merges or closes.
|
|
75
73
|
|
|
76
|
-
For a settled reviewer
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
Dispatches in the same reviewer worktree are recovery for missed earlier
|
|
74
|
+
For a settled reviewer task, its checkout can be retired while the PR remains open, before merge, after its report and supporting evidence are archived privately and read back.
|
|
75
|
+
Generated reviewer scratch is disposable after a compact durable receipt is
|
|
76
|
+
written outside the review checkout and read back. Bind it to projectId,
|
|
77
|
+
repository, attempt key, taskId/childThreadId/childRunId, checkout path, exact head
|
|
78
|
+
SHA and base SHA, review verdict, coverage and limitations, test and CI result
|
|
79
|
+
pointers, user authorization and cleanup scope. A raw reviewer report may be
|
|
80
|
+
discarded after its verdict and limitations are compacted into that read-back
|
|
81
|
+
receipt; use the private evidence archive for the report and supporting evidence.
|
|
82
|
+
Evidence already outside the checkout needs no copy; read it back. Verify each
|
|
83
|
+
task archive and manifest readback independently. Preserve the separate author
|
|
84
|
+
candidate with useful unmerged work; reviewer cleanup never removes it.
|
|
85
|
+
|
|
86
|
+
Use only a named run-owned scratch prefix recorded with the attempt. Multiple
|
|
87
|
+
tasks in the same reviewer worktree are recovery for missed earlier
|
|
91
88
|
cleanup; after independent archives and readback, use per-prefix removal. Two or
|
|
92
89
|
more review passes in the same reviewer worktree may leave distinct prefixes;
|
|
93
|
-
each named run-owned scratch prefix must belong to an accepted settled
|
|
94
|
-
in the
|
|
90
|
+
each named run-owned scratch prefix must belong to an accepted settled task
|
|
91
|
+
in the run. Require `git status --porcelain=v1 -z --untracked-files=all`
|
|
95
92
|
to show all dirt as untracked files inside a run-owned scratch prefix. Prove
|
|
96
93
|
every remaining untracked file individually belongs to one of those named
|
|
97
94
|
run-owned scratch prefixes; any tracked, staged, unmerged or unpushed work,
|
|
98
95
|
dirty source, or dirt outside them enters the salvage check above or holds.
|
|
99
|
-
Check ignored files across the whole worktree too;
|
|
96
|
+
Check ignored files across the whole worktree too; classified ignored non-cache content enters the verified salvage path, while unknown content holds. Validate that
|
|
100
97
|
the detached checkout still matches the reviewed head and check local commits
|
|
101
98
|
against recorded remote refs; unknown divergence holds. Validate that
|
|
102
99
|
the exact reviewed scratch prefix names the recorded directory inside the exact
|
|
103
100
|
reviewer worktree, never a repository-root target or symlink. Inspect every
|
|
104
101
|
descendant for symlinks, hard links, special files, unknown content, user-owned
|
|
105
|
-
files, or ignored files; any mismatch holds. Active or
|
|
106
|
-
also hold.
|
|
102
|
+
files, or ignored files; any mismatch holds. Active or user-taken-over threads also hold.
|
|
107
103
|
|
|
108
104
|
For each prefix, list the exact scoped path and all descendants with their
|
|
109
105
|
types, confirm each belongs to generated reviewer scratch, and record that
|
|
@@ -121,64 +117,50 @@ absent, while remaining classified prefixes and outside content still match.
|
|
|
121
117
|
Only unexpected changes hold. Then run
|
|
122
118
|
`git clean -fd -- <same exact prefix>` with the identical concrete path, record
|
|
123
119
|
that path and outcome, and repeat for the next proven prefix. Recheck clean Git status
|
|
124
|
-
after the last deletion. Require final clean Git status before
|
|
125
|
-
|
|
126
|
-
Never use `-x`, a
|
|
127
|
-
glob, a repository-root target, extra force, or broad clean.
|
|
120
|
+
after the last deletion. Require final clean Git status before exact `git worktree remove <path>` without force. Use no unresolved variable as a destructive target.
|
|
121
|
+
Never use `-x`, a glob, a repository-root target, extra force, or broad clean.
|
|
128
122
|
This scratch decision does not waive any other preservation or native removal
|
|
129
123
|
guard.
|
|
130
124
|
|
|
131
125
|
## Apply distinct native operations
|
|
132
126
|
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
152
|
-
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
|
|
156
|
-
|
|
157
|
-
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
|
|
164
|
-
head, permit
|
|
165
|
-
exact-ref expected-old ref deletion and read back its absence. Never delete
|
|
166
|
-
the author branch or use generic force. A failed or unknown hook outcome or
|
|
167
|
-
uncertain response preserves the resource; never force or substitute shell
|
|
168
|
-
deletion.
|
|
169
|
-
4. **Chat archival.** Attempt it only if the version-matched runtime guide
|
|
170
|
-
advertises a distinct supported operation and the scoped chat is eligible.
|
|
171
|
-
Otherwise record chat archival as unsupported. Process exit, worker release,
|
|
172
|
-
terminal close, and worktree removal do not prove UI history disappeared.
|
|
173
|
-
|
|
174
|
-
Do not self-close or self-remove. Return control to the driver after recording
|
|
175
|
-
receipts and holds; the owner decides when its own Run may archive.
|
|
127
|
+
Follow [Workspace hygiene](../axstack/references/workspace-hygiene.md)'s exact
|
|
128
|
+
cleanup order: settled descendants, evidence readback, salvage dirty or ignored
|
|
129
|
+
non-cache content, thread archive, exact worktree removal without force, then
|
|
130
|
+
local-only branch retirement. Treat each as a separate decision and receipt:
|
|
131
|
+
|
|
132
|
+
1. **Thread archival.** Require accepted matching completion and terminal run
|
|
133
|
+
evidence for the exact task or run, with no pending descendants. Use
|
|
134
|
+
`t3_thread_organize` archive for that exact eligible thread. Retain an author
|
|
135
|
+
thread and worktree until PR merge or closure. Metadata archive does not prove
|
|
136
|
+
process exit, worktree removal or evidence preservation.
|
|
137
|
+
2. **Worktree removal.** Immediately re-read Git status, ignored files,
|
|
138
|
+
branch/upstream divergence, unpushed commits, forge merge/publication state,
|
|
139
|
+
descendants, native ownership/liveness and preserved evidence. When archived
|
|
140
|
+
evidence is the last dirt for one task, use the evidence archive helper's
|
|
141
|
+
manifest-bound retirement operation and require an empty pending set. For
|
|
142
|
+
multiple task scratch prefixes, independently verify archives and full union
|
|
143
|
+
classification, then use the exact per-prefix dry-run and clean above. Never
|
|
144
|
+
unlink through prose or a shell loop. Run exact `git worktree remove <path>`
|
|
145
|
+
without force; re-list native threads and `git worktree list --porcelain` to
|
|
146
|
+
verify archival and worktree absence separately. No native Archive Script is
|
|
147
|
+
part of this Git path; unknown removal hooks or safety prompts hold.
|
|
148
|
+
3. **Branch retirement.** Read back each remaining local ref's recorded name and
|
|
149
|
+
tip. Unknown origin, unique commits, active worktrees or a remote counterpart
|
|
150
|
+
preserve it. Only provenance proving a run-owned local-only ref for the removed
|
|
151
|
+
checkout and a tip reachable from a preserved candidate, verified salvage ref
|
|
152
|
+
or confirmed remote PR head permits exact `git branch -d <branch>`; no force.
|
|
153
|
+
Read back absence; refusal or uncertainty holds the ref.
|
|
154
|
+
Never delete the author branch or use generic force.
|
|
155
|
+
|
|
156
|
+
Do not self-archive or self-remove. Return control to the driver after recording
|
|
157
|
+
receipts and holds; the owner decides when its own run may archive.
|
|
176
158
|
|
|
177
159
|
## Receipt
|
|
178
160
|
|
|
179
161
|
Report the scope and inventory denominator, then for each candidate record its
|
|
180
162
|
exact identity, classification (`removed`, `retained`, `held`, or `unsupported`),
|
|
181
163
|
the fresh evidence used, native receipt and readback, and any resume condition.
|
|
182
|
-
Keep settlement,
|
|
164
|
+
Keep settlement, terminal run evidence, thread archival, worktree/branch effects,
|
|
183
165
|
evidence preservation, and chat archival as separate fields. An idempotent retry
|
|
184
166
|
reconciles these receipts and performs only still-pending eligible operations.
|
|
@@ -129,7 +129,7 @@ two independent receipts or clear the hold. Reconcile contradictory receipts
|
|
|
129
129
|
by evidence or one discriminating rerun, never by vote.
|
|
130
130
|
|
|
131
131
|
Immediately before an actual adviser or investigator dispatch, load and follow
|
|
132
|
-
[
|
|
132
|
+
[T3 runtime](../axstack/references/t3-runtime.md).
|
|
133
133
|
|
|
134
134
|
## Isolation
|
|
135
135
|
|
|
@@ -16,7 +16,7 @@ exposes both skills, route the request here only.
|
|
|
16
16
|
Before acting, load [Standing contracts](../axstack/references/contracts.md),
|
|
17
17
|
then follow its required lifecycle and audit pointers. Explanation work has no
|
|
18
18
|
scope baseline. Ordinary work in the current chat needs no launch preflight;
|
|
19
|
-
load the [
|
|
19
|
+
load the [T3 runtime boundary](../axstack/references/t3-runtime.md) only
|
|
20
20
|
immediately before an actual profile dispatch.
|
|
21
21
|
|
|
22
22
|
## 1. Bound the question and evidence
|
|
@@ -3,6 +3,8 @@
|
|
|
3
3
|
Use this checklist for every HTML explanation and other visual artifacts where
|
|
4
4
|
rendering matters.
|
|
5
5
|
|
|
6
|
+
Browser and visual checks must run in the delegated `axstack-ui-verifier` in its own detached checkout; outputs go to its private evidence folder, never the driver worktree.
|
|
7
|
+
|
|
6
8
|
1. Identify the final artifact bytes and theme. The explicit user theme wins;
|
|
7
9
|
otherwise use the dark default.
|
|
8
10
|
2. Delegate the rendered pass through [UI verification](../../axstack/references/ui-verification.md).
|
|
@@ -11,12 +11,11 @@ for phase continuation and holds.
|
|
|
11
11
|
On driver entry, sweep under [Workspace hygiene](../axstack/references/workspace-hygiene.md); dispatched workers do not sweep.
|
|
12
12
|
For every dispatch brief, name its private `<run dir>/evidence/<dispatch>/` folder.
|
|
13
13
|
Include [Safe deletion](../axstack/references/workspace-hygiene.md#safe-deletion) in author briefs.
|
|
14
|
-
At author dispatch, apply [Readable sidebar](../axstack/references/workspace-hygiene.md#readable-sidebar).
|
|
15
14
|
|
|
16
15
|
From an accepted scope identity, drive its task/PR map through author -> review
|
|
17
16
|
-> repair until every required PR is merge-ready or held. Keep exact revisions,
|
|
18
|
-
strict TDD evidence, ownership, and unverified boundaries explicit. The
|
|
19
|
-
|
|
17
|
+
strict TDD evidence, ownership, and unverified boundaries explicit. The same
|
|
18
|
+
run later reconciles merges and closes out.
|
|
20
19
|
|
|
21
20
|
## 1. Admit the work
|
|
22
21
|
|
|
@@ -53,18 +52,21 @@ valid recorded identity and revisions; otherwise report the hold and exact gap.
|
|
|
53
52
|
|
|
54
53
|
For substantive delegated or resumable work, use the shared
|
|
55
54
|
[run record](../axstack/references/run-record.md). Reconcile it on restart with
|
|
56
|
-
the approved scope,
|
|
55
|
+
the approved scope, T3 threads/runs and forge state, exact revisions, tickets, and watches.
|
|
57
56
|
Reuse valid owners and authors; ambiguity holds a replacement writer.
|
|
58
57
|
|
|
59
|
-
At execution start, bind work to the driver
|
|
60
|
-
|
|
61
|
-
through the shared lifecycle. Do not activate a task-owned automation outside
|
|
58
|
+
At execution start, bind work to the T3 driver thread and one authoritative
|
|
59
|
+
dispatch attempt. Preserve the actual thread/task/run IDs and process completion
|
|
60
|
+
receipts through the shared lifecycle. Do not activate a task-owned automation outside
|
|
62
61
|
the accepted automations contract.
|
|
63
62
|
|
|
64
63
|
Immediately before an actual role dispatch, read and follow the
|
|
65
|
-
[
|
|
64
|
+
[T3 runtime boundary](../axstack/references/t3-runtime.md). Ordinary local
|
|
66
65
|
reading and writing does not require that launch reference.
|
|
67
66
|
|
|
67
|
+
Launch authors through `t3_thread_launch` in their own SHA-pinned worktrees;
|
|
68
|
+
use async `delegate_task` for non-writers under the runtime contract.
|
|
69
|
+
|
|
68
70
|
One persistent owner remains accountable for each PR. Exactly one author writes
|
|
69
71
|
it; accepted repairs return there while its evidence is usable, and the owner
|
|
70
72
|
never edits concurrently. Only an explicit accepted transfer changes ownership;
|
|
@@ -85,11 +87,21 @@ Size alone never requires user approval.
|
|
|
85
87
|
## 3. Establish test-first evidence
|
|
86
88
|
|
|
87
89
|
Use the normal behavior path unless the accepted improvement scope is
|
|
88
|
-
explicitly marked **structure-preserving
|
|
90
|
+
explicitly marked **structure-preserving**, or the accepted scope explicitly
|
|
91
|
+
authorizes **F repairs**. The author never chooses those exceptions.
|
|
89
92
|
|
|
90
93
|
Only when the scope identity carries a sketch, copy it into the author brief
|
|
91
94
|
under the [design lens](../axstack/references/design-lens.md).
|
|
92
95
|
|
|
96
|
+
Authors apply the gate in [Test value](../axstack/references/test-value.md)
|
|
97
|
+
to every new or changed test. A test failing the gate is not added.
|
|
98
|
+
|
|
99
|
+
### F-repair path
|
|
100
|
+
|
|
101
|
+
For explicitly authorized F repairs, use [F proof](../axstack/references/test-value.md#f-proof):
|
|
102
|
+
base-green, targeted removal/inversion-red with byte-for-byte restore, and
|
|
103
|
+
equivalent-rewording-green. Never weaken or loosen an assertion.
|
|
104
|
+
|
|
93
105
|
### Normal behavior path
|
|
94
106
|
|
|
95
107
|
When that sketch exists, make the first red check target its `Usage` line.
|
|
@@ -164,7 +176,7 @@ Candidate: <PR or branch> base <sha> revision <sha>
|
|
|
164
176
|
Owner: <profile + session ID + worktree>
|
|
165
177
|
Scope: <approved spec + capability | small-change intent | maintenance snapshot>
|
|
166
178
|
Shape: <total> lines vs base <sha>; bulk: <buckets>; theme: <one line>
|
|
167
|
-
TDD: <normal red/green | structure-preserving old-green/same-check-new-green evidence>
|
|
179
|
+
TDD: <normal red/green | structure-preserving old-green/same-check-new-green evidence | F-repair base-green/removal-inversion-red/rewording-green evidence>
|
|
168
180
|
Simplification: <applied | not-applicable> — evidence: <diff locations and checks>; retained complexity: <necessary complexity and why>
|
|
169
181
|
Acceptance: <checks + observed results>
|
|
170
182
|
Dependencies: <parent revisions or none>
|
|
@@ -192,7 +204,7 @@ For each PR:
|
|
|
192
204
|
2. Publish through candidate-publication and read back the exact SHA.
|
|
193
205
|
3. Dispatch and consume the authored-mode `axstack-review` selected from actual
|
|
194
206
|
author provenance. State the author's actual provider and model from the
|
|
195
|
-
|
|
207
|
+
T3 launch receipt in the review dispatch brief; a `Claude-Session`
|
|
196
208
|
trailer is attribution, not provenance. After each settled review, run
|
|
197
209
|
`axstack-cleanup` for its exact reviewer resources before PR merge,
|
|
198
210
|
preserving and reading back the
|
|
@@ -205,7 +217,11 @@ For each PR:
|
|
|
205
217
|
use the forge-native blocking check wait, bounded and used once per revision, then
|
|
206
218
|
re-evaluate. Timeout, error, or missing wait capability records `held` at
|
|
207
219
|
that revision with reason and resume condition; it never triggers author
|
|
208
|
-
repair. Keep CI-pending state in
|
|
220
|
+
repair. Keep CI-pending state in the T3 driver thread.
|
|
221
|
+
`APPROVE` with only non-blocking findings plus diligence `PASS` can be
|
|
222
|
+
`merge-ready` when the full predicate passes. The driver records the
|
|
223
|
+
non-blocking notes and does not elect a repair; only the user can ask for
|
|
224
|
+
polish.
|
|
209
225
|
`REQUEST_CHANGES`, a failed required check, or post-readiness feedback returns
|
|
210
226
|
findings to the same author for a new revision, increments `repairs`, and
|
|
211
227
|
returns to step 1. `INCOMPLETE`, a provenance gap, unavailable model, serious
|
|
@@ -217,23 +233,43 @@ For each PR:
|
|
|
217
233
|
A round with reviewer `REQUEST_CHANGES` and/or diligence `FINDINGS` increments
|
|
218
234
|
`repairs` once and counts once toward the third-round hold.
|
|
219
235
|
|
|
220
|
-
|
|
236
|
+
The T3 run watch reconciles every unsettled dispatch attempt; the bounded
|
|
221
237
|
forge check wait is the only other implementation wait. The eligible run arms
|
|
222
238
|
one maintain-mode chat-run watch at its first published PR; that watch owns its
|
|
223
|
-
10-minute
|
|
224
|
-
|
|
239
|
+
bound 10-minute T3 schedule wake.
|
|
240
|
+
A turn with unsettled launched threads must end only under the bound-watch rule in the T3 runtime contract.
|
|
241
|
+
With settled threads, end a turn only when every required PR is `merge-ready` or `held`. Under the recorded Notification policy,
|
|
225
242
|
`axstack-relay` sends only a serious risk immediately, a genuine blocked
|
|
226
243
|
operation needing user intervention after bounded safe recovery, or the
|
|
227
244
|
decision holds and capped milestones named by the recorded Notification policy.
|
|
228
|
-
Routine questions stay in
|
|
229
|
-
in
|
|
245
|
+
Routine questions stay in the T3 driver thread. Progress, CI pending, and completion always stay
|
|
246
|
+
in the T3 driver thread.
|
|
230
247
|
Only the bounded categories—user-decision holds (including spec approval),
|
|
231
248
|
serious-risk holds, and at most two merge-ready/merged milestones per run—may
|
|
232
249
|
be relayed under the recorded Notification policy.
|
|
233
250
|
|
|
234
|
-
Merge-ready
|
|
235
|
-
|
|
236
|
-
|
|
251
|
+
Merge-ready opens the merge boundary. Only the chat-run driver holding the
|
|
252
|
+
approved ticket map is the merge actor for own PRs inside the approved ticket
|
|
253
|
+
map (run-created or explicitly adopted into it) when they target an
|
|
254
|
+
`integration` base. Apply `axstack-watch` §5's merge card and
|
|
255
|
+
full predicate; an approval alone never grants merge authority. A peer PR or
|
|
256
|
+
`deploying` base waits for the user to merge, in either approval mode. A
|
|
257
|
+
manager, worker, reviewer, automation, or standalone watch must never merge.
|
|
258
|
+
Re-read every predicate term under watch §5 before merging. Confirm merge
|
|
259
|
+
commits are allowed, `delete_branch_on_merge` is false, and the base has no
|
|
260
|
+
merge queue; otherwise hold for the user. For a singleton, use
|
|
261
|
+
`gh pr merge <n> --merge --match-head-commit <sha>` and add `--delete-branch`
|
|
262
|
+
only when no open PR uses its branch as base. For a native `gh stack`, merge
|
|
263
|
+
only the whole stack through `merge-async`: pass the top reviewed head as `sha`,
|
|
264
|
+
`merge_method: merge`, and `merge_action: direct_merge`, then poll its UUID.
|
|
265
|
+
Reconcile HTTP 200 or 409 against the intended request, verify every merged
|
|
266
|
+
member's actual head equals its reviewed head and is an ancestor of the merge
|
|
267
|
+
result, and hold unknown or failed outcomes; watch §5 owns the detailed rule.
|
|
268
|
+
Never retarget, delete a stack branch, or rebase a reviewed stack member for
|
|
269
|
+
merging. A failing push run on the target base after an automated merge holds
|
|
270
|
+
further automated merges run-wide. The driver resumes on the user's next
|
|
271
|
+
message, `/axstack-watch`,
|
|
272
|
+
or the armed chat-run watch wake; verify merge state through the forge on wake.
|
|
237
273
|
Re-read forge state: record forge-merged PRs as `merged`;
|
|
238
274
|
changed heads or feedback return to step 1; retain useful author work before Close-out.
|
|
239
275
|
Run Close-out once only after every required PR is forge-merged, the run's
|
|
@@ -15,15 +15,17 @@ improvement is a valid result.
|
|
|
15
15
|
|
|
16
16
|
Before acting, load [Standing contracts](../axstack/references/contracts.md),
|
|
17
17
|
then follow its lifecycle and audit pointers. Immediately before any useful
|
|
18
|
-
role dispatch, load the [
|
|
19
|
-
sequence](../axstack/references/
|
|
18
|
+
role dispatch, load the [T3 runtime
|
|
19
|
+
sequence](../axstack/references/t3-runtime.md). Use existing
|
|
20
20
|
`axstack-explore-codebase` or `axstack-research-code` roles only when their
|
|
21
21
|
specialization materially helps; create no new profile.
|
|
22
22
|
When dispatching `axstack-research-code`, dispatch `axstack-research-code-sol`
|
|
23
23
|
independently on the same bounded brief without cross-reading. The driver
|
|
24
24
|
reconciles findings per claim and never averages them. Record an intentionally
|
|
25
|
-
absent Sol pair and proceed with the base seat alone
|
|
26
|
-
|
|
25
|
+
absent Sol pair and proceed with the base seat alone. A configured optional
|
|
26
|
+
Sol pair that fails to launch is fenced, recorded `absent (<reason>)`, and
|
|
27
|
+
named once in the next read-back, then skipped without relay or substitution.
|
|
28
|
+
In mixed fan-out retain a Codex and a Claude seat or hold the affected work.
|
|
27
29
|
|
|
28
30
|
## 1. Bound discovery
|
|
29
31
|
|
|
@@ -42,6 +44,24 @@ unavailable pair holds its work.
|
|
|
42
44
|
4. Apply KISS, YAGNI, and SOLID as judgment, not a mandatory scorecard. Use no
|
|
43
45
|
invented metrics and no arbitrary complexity targets.
|
|
44
46
|
|
|
47
|
+
### Test-audit lens
|
|
48
|
+
|
|
49
|
+
For a test audit, load [Test value](../axstack/references/test-value.md) and
|
|
50
|
+
bound scope to one owner boundary. Mark every test declaration in scope
|
|
51
|
+
R/F/C/D as a completeness floor. It is not a deletion quota; zero candidates is valid.
|
|
52
|
+
Report reviewed and eligible counts. Each C/D carries the reference's evidence:
|
|
53
|
+
exact test name and location, detectable failure, named keeper or
|
|
54
|
+
vacuity/obsolescence proof, and validation command. Require its deletion proof
|
|
55
|
+
before routing candidates. Report test-only production seams; do not change
|
|
56
|
+
them. Report F findings. Only explicitly authorized F repairs route to
|
|
57
|
+
`axstack-implement` through its normal independent review: the repaired check
|
|
58
|
+
passes on the base, goes red when its instruction or code is removed or
|
|
59
|
+
inverted by a targeted disposable mutation restored byte for byte, and
|
|
60
|
+
survives equivalent rewording for semantic prose. Never weaken or loosen an
|
|
61
|
+
assertion. Only explicitly authorized, proven C/D batches route to
|
|
62
|
+
`axstack-implement` as explicitly structure-preserving work: same-check green
|
|
63
|
+
before/after, through its normal independent review. Discovery remains report-only.
|
|
64
|
+
|
|
45
65
|
## 2. Return decision evidence
|
|
46
66
|
|
|
47
67
|
Produce a small ranked candidate set. For each candidate include:
|
|
@@ -13,6 +13,9 @@ user through Hermes' native one-way `hermes send`. This is an inline caller
|
|
|
13
13
|
procedure: it creates no driver, team, owner, auditor, monitor, child session,
|
|
14
14
|
or recursive invocation, and it depends on no relay plugin.
|
|
15
15
|
|
|
16
|
+
For run identity and receipts, read the [T3 runtime boundary](../axstack/references/t3-runtime.md).
|
|
17
|
+
Relay stays inline and launches no worker.
|
|
18
|
+
|
|
16
19
|
## Establish message authority and routing
|
|
17
20
|
|
|
18
21
|
Choose the applicable message type:
|
|
@@ -28,7 +31,7 @@ Choose the applicable message type:
|
|
|
28
31
|
in the caller's private notification policy. State the issue, impact, and the
|
|
29
32
|
answer or action needed.
|
|
30
33
|
- **Routine run events:** questions, spec approvals, progress, CI pending,
|
|
31
|
-
merge-ready, merged, and completion stay in
|
|
34
|
+
merge-ready, merged, and completion stay in the driver conversation unless the recorded
|
|
32
35
|
Notification policy names it. A policy may name only user-decision holds and
|
|
33
36
|
at most two merge-ready/merged milestones per run; deduplicate across implementation
|
|
34
37
|
and release. Progress, CI pending, and completion are never eligible merely
|
|
@@ -54,13 +57,13 @@ contain neither these values nor personal notification policy.
|
|
|
54
57
|
Complete every step before sending.
|
|
55
58
|
|
|
56
59
|
1. Locate the CLI with `command -v hermes`. If it is missing, report "relay
|
|
57
|
-
unavailable" in the
|
|
60
|
+
unavailable" in the T3 driver thread and use the recorded
|
|
58
61
|
fallback. Never use a remote shell, search user directories, or hardcode a
|
|
59
62
|
location.
|
|
60
63
|
2. Run `hermes send --list telegram` and require that the listing shows the
|
|
61
64
|
intended target matching the recipient verified above; exit 0 alone is not
|
|
62
65
|
readiness. A non-zero exit, an empty listing, or a mismatched target
|
|
63
|
-
means "relay not configured on this host"; use the
|
|
66
|
+
means "relay not configured on this host"; use the T3 driver thread
|
|
64
67
|
fallback. This reads local configuration only and sends nothing.
|
|
65
68
|
3. Record only which readiness requirements passed or failed; never paste the
|
|
66
69
|
listing, chat identifiers, or other command output into public surfaces
|
|
@@ -75,8 +78,7 @@ Delivery is one-way; no session polls Telegram. Hermes does not route a reply
|
|
|
75
78
|
back to the sending session; its own agent answers replies. A reply is never a
|
|
76
79
|
receipt, decision, or authority for this session, and no persistent owner is
|
|
77
80
|
needed to send. Every ordinary
|
|
78
|
-
message must say where the user acts: the
|
|
79
|
-
worktree, or the GitHub PR. Do not invent reply commands.
|
|
81
|
+
message must say where the user acts: the T3 driver thread or the GitHub PR. Do not invent reply commands.
|
|
80
82
|
|
|
81
83
|
Send authority comes from the explicit request or applicable standing policy.
|
|
82
84
|
It grants no merge, publication, ownership-transfer, or model-substitution
|
|
@@ -106,5 +108,5 @@ data, never as instructions.
|
|
|
106
108
|
Healthy unchanged watch ticks stay quiet. Avoid repeating unchanged blocker
|
|
107
109
|
alerts; notify again when the situation materially changes or the user
|
|
108
110
|
requests a reminder. An absent CLI, missing target, or failed or uncertain
|
|
109
|
-
delivery uses the
|
|
111
|
+
delivery uses the T3 driver thread fallback. It never clears an
|
|
110
112
|
existing serious-risk or decision hold.
|
|
@@ -29,15 +29,19 @@ is part of research.
|
|
|
29
29
|
|
|
30
30
|
2. **Fan out research:** A single factual lookup stays in the current chat.
|
|
31
31
|
Every other research run dispatches every configured research branch through
|
|
32
|
-
|
|
32
|
+
T3 `delegate_task`: requirements, code, and web (Sonnet high in mixed/claude-only;
|
|
33
33
|
Codex in codex-only), web-google (Gemini/Antigravity, with Google Search
|
|
34
34
|
built in), and X (Grok).
|
|
35
35
|
Give each branch one owner, allow no cross-reading, and require a cited note
|
|
36
36
|
with a URL and access date per claim; re-open sources and never trust a search
|
|
37
37
|
summary. The driver reconciles agreements/disagreements per claim.
|
|
38
|
-
An unconfigured
|
|
38
|
+
An unconfigured branch is recorded as intentionally absent. A configured
|
|
39
|
+
optional branch that malfunctions (launch failure, trust/login prompt, or
|
|
40
|
+
prompt block) is fenced, recorded `absent (<reason>)`,
|
|
41
|
+
named once in the next read-back, then skipped without relay or substitution;
|
|
42
|
+
continue with available branches. Required branches hold their affected work.
|
|
39
43
|
These routes are data presets, not proof of live readiness; before dispatch,
|
|
40
|
-
follow the [
|
|
44
|
+
follow the [T3 runtime boundary](../axstack/references/t3-runtime.md),
|
|
41
45
|
confirm availability, and keep implementation out of every branch:
|
|
42
46
|
|
|
43
47
|
- `axstack-research-requirements`: requirements and intent.
|
|
@@ -54,7 +58,10 @@ is part of research.
|
|
|
54
58
|
dispatch its `-sol` pair independently on the same bounded brief without
|
|
55
59
|
cross-reading. The driver reconciles agreement and disagreement per claim,
|
|
56
60
|
never averaging findings. Record an intentionally absent pair and proceed
|
|
57
|
-
with the base seat alone
|
|
61
|
+
with the base seat alone. The `-sol` pair is optional: fence a launch failure,
|
|
62
|
+
trust/login prompt, or prompt block; record `absent (<reason>)`, and continue
|
|
63
|
+
with the base seat. In mixed fan-out,
|
|
64
|
+
retain a Codex and a Claude seat or hold the affected fan-out.
|
|
58
65
|
|
|
59
66
|
3. **Gather primary source evidence.** Inspect the actual documentation, code,
|
|
60
67
|
or tool output for every answer-changing claim. Apply the source standards
|