axstack 0.20.30 → 0.21.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (48) hide show
  1. package/README.md +24 -23
  2. package/bin/axstack.js +18 -5
  3. package/docs/installation.md +101 -46
  4. package/docs/workflows.md +179 -117
  5. package/package.json +3 -3
  6. package/profiles/presets/claude-only.json +46 -46
  7. package/profiles/presets/codex-only.json +50 -50
  8. package/profiles/presets/mixed.json +59 -59
  9. package/skills/axstack/references/automations.md +127 -137
  10. package/skills/axstack/references/autopilot.md +121 -0
  11. package/skills/axstack/references/candidate-publication.md +13 -8
  12. package/skills/axstack/references/contracts.md +10 -4
  13. package/skills/axstack/references/diligence.md +3 -1
  14. package/skills/axstack/references/evidence-archive.md +38 -33
  15. package/skills/axstack/references/lifecycle.md +64 -50
  16. package/skills/axstack/references/review-manager-prompt.md +13 -11
  17. package/skills/axstack/references/role-roster.md +19 -9
  18. package/skills/axstack/references/routing.md +33 -18
  19. package/skills/axstack/references/run-record.md +36 -15
  20. package/skills/axstack/references/t3-runtime.md +234 -0
  21. package/skills/axstack/references/test-audit-weekly.md +62 -0
  22. package/skills/axstack/references/test-value.md +120 -0
  23. package/skills/axstack/references/ui-verification.md +5 -1
  24. package/skills/axstack/references/workspace-hygiene.md +102 -156
  25. package/skills/axstack/scripts/pr-digest.js +120 -0
  26. package/skills/axstack/scripts/resolve-models.js +102 -0
  27. package/skills/axstack-align/SKILL.md +25 -11
  28. package/skills/axstack-audit/SKILL.md +22 -5
  29. package/skills/axstack-audit/references/record.md +1 -1
  30. package/skills/axstack-cleanup/SKILL.md +69 -87
  31. package/skills/axstack-debug/SKILL.md +1 -1
  32. package/skills/axstack-explain/SKILL.md +1 -1
  33. package/skills/axstack-explain/references/visual-qa.md +2 -0
  34. package/skills/axstack-implement/SKILL.md +76 -26
  35. package/skills/axstack-improve/SKILL.md +24 -4
  36. package/skills/axstack-relay/SKILL.md +16 -7
  37. package/skills/axstack-research/SKILL.md +11 -4
  38. package/skills/axstack-review/SKILL.md +42 -32
  39. package/skills/axstack-spec/SKILL.md +23 -14
  40. package/skills/axstack-tickets/SKILL.md +13 -11
  41. package/skills/axstack-watch/SKILL.md +117 -34
  42. package/skills/axstack-watch/references/watch-runtime.md +61 -69
  43. package/src/capabilities.js +33 -69
  44. package/src/installer.js +9 -1
  45. package/src/instructions.js +9 -4
  46. package/src/roles.js +38 -10
  47. package/skills/axstack/references/orca-runtime.md +0 -183
  48. package/skills/axstack/scripts/trust-path.js +0 -123
@@ -1,170 +1,116 @@
1
- # Workspace hygiene for driver-owned Orca runs
1
+ # Workspace hygiene for driver-owned T3 runs
2
2
 
3
- This is a prompt contract for drivers, not a cleanup daemon or new Orca
4
- protocol. Use the version-matched Orca guides for native operations. Dispatched
5
- workers never sweep or remove another session. A scheduled pass with recorded
6
- cleanup authority acts as its lane's driver for the sweep; read-only observers
7
- only report leftovers. Record each decision and native readback in the private
8
- run record; uncertain ownership, liveness, or evidence
9
- holds only the affected resource.
3
+ This is a prompt contract, not a cleanup daemon. Follow [T3 runtime](t3-runtime.md)
4
+ for native identities, terminal run evidence and schema. Dispatched workers never
5
+ sweep or remove another session. A scheduled pass with recorded cleanup authority
6
+ acts as its lane's driver; read-only observers only report leftovers. Uncertain
7
+ ownership, liveness or evidence holds only the affected resource.
8
+ Record each decision and native readback in the private run record.
10
9
 
11
10
  ## Safe deletion
12
11
 
13
- Every shell deletion targets a literal absolute path or a `${VAR:?}`-guarded
14
- expansion (for example, `rm -rf -- "${EV:?}/mut"`), only inside the worker's
15
- own evidence folder, `TMPDIR`, or worktree. Never use a bare `$VAR`, a glob on a
16
- variable, `/`, `HOME`, or a shared root as a deletion target. Prefer
17
- `git clean -- <exact prefix>` or tool-native cleanup. A safety prompt that
12
+ Before use, commands must scope `TMPDIR` to an owned 0700 directory under the system temp directory, never under `$HOME`, named from the dispatch key and recorded in the receipt.
13
+ Validate its real path, absence of symlinks and ownership before use and cleanup; remove it afterwards by literal absolute path.
14
+ Evidence files still go to the private `<run>/evidence/<key>/` folder.
15
+
16
+ Every shell deletion targets a literal absolute path or a `${VAR:?}`-guarded expansion, only inside the worker's own evidence folder, `TMPDIR`, or worktree.
17
+ For example, `rm -rf -- "${EV:?}/mut"` requires a validated owned evidence path.
18
+ Never use a bare `$VAR`, a glob on a variable, `/`, `HOME`, or a shared root as a deletion target.
19
+ Prefer `git clean -- <exact prefix>` or tool-native cleanup. A safety prompt that
18
20
  still appears is a hold; agents do not answer it.
19
21
 
22
+ ## Project preflight
23
+
24
+ `worktreeCleanup` must be `off` for every Axstack project.
25
+ Preflight reads back `worktreeCleanup` via `t3_project_read` where exposed; otherwise record a limitation pointing to the documented installation setup step.
26
+ Automatic deletion cannot replace evidence readback or salvage. Do not change
27
+ project settings without recorded host-mutation authority.
28
+
20
29
  ## Settlement
21
30
 
22
- At intermediate completion, once a worker or reviewer Dispatch is accepted,
23
- the driver releases it natively, confirms closure from a fresh native terminal
24
- list, then removes its worktree, descendants first, after the preservation
25
- checks below. Keep the author worktree and session until the PR merges or closes
26
- so repairs return to the same author. Release, terminal closure, worktree removal, and branch
27
- retirement each need their own receipt.
28
-
29
- At final settlement, no eligible non-driver session or worktree remains,
30
- including the run's worktrees in other repositories. Report each remaining
31
- resource as a hold with its reason. Recorded native ownership by Run, Task,
32
- and Dispatch decides; parent/child display lineage does not. The creator closes
33
- what it created at ordinary settlement; the cross-run sweep below may retire
34
- its settled leftovers. A finite scheduled pass closes only its own exact terminal as
35
- its final action. Manual chats, automation dedicated workspaces, and genuine
36
- `user_takeover` sessions are never removed by ordinary settlement or sweep.
37
- Deleting a session means closing its terminal; agent chat history is not deleted.
38
-
39
- If native exact terminal close returns `runtime_error` for a provably finished
40
- agent, send `/quit` + Enter to that exact terminal, wait about 5 seconds, then
41
- send `exit` + Enter. Confirm it left a fresh native terminal list. Never use
42
- this fallback for a working, user-taken-over, or unclear agent; never use
43
- `--all` or a name selector. A finite scheduled pass whose own close fails
44
- leaves its terminal for the next pass, without treating that expected failure
45
- as a hold. At pass start, clear only provably finished predecessor terminals
46
- of the same automation in its dedicated workspace by this exact-handle path.
47
-
48
- ## Owned automation retirement
49
-
50
- At Close-out, reconcile the run record with native inventory: the recorded
51
- owning Run must have created the exact automation IDs selected for retirement.
52
- The sweeping pass uses that recorded owning Run, not its own Run, for cross-run
53
- watches. A cross-run sweep may retire a disabled per-run watch only when its
54
- owning run is closed or every watched PR is merged or closed. Uncertain recorded
55
- ownership, watch state, or disable result holds the affected automation.
56
-
57
- For each eligible automation, disable it and verify native readback before
58
- `orca automations remove <id>` on its exact ID, and verify absence by native
59
- readback. The observer only disables and reports; the driver removes its own
60
- run's automation. Only a retired owned per-run watch's workspace may be removed
61
- under these guards. Never remove the durable review manager and its dedicated
62
- workspace or a user-created automation and its dedicated workspace.
63
-
64
- Remove that dedicated workspace only after confirming its exact ownership, no
65
- live terminal, a clean worktree, and a head on the remote. Recheck ownership and
66
- liveness immediately before workspace removal. If dirty or unpublished, use the
67
- salvage and bundle verification below before removal; failed verification or
68
- uncertain liveness is a hold. The current scheduled pass cannot remove its own
69
- workspace while its terminal is live; the driver or a later sweep finishes that
70
- step. Record separate automation and workspace receipts.
31
+ Cleanup order is: settled descendants -> evidence readback -> salvage dirty or ignored non-cache content -> t3_thread_organize archive -> exact git worktree remove without force -> git branch -d for local-only branches.
32
+ Require terminal run evidence before `t3_thread_organize` settle or archive.
33
+ These metadata actions do not remove worktrees. Accept matching delegated task
34
+ or launched run completion under the runtime contract before settlement.
35
+ Keep the author worktree and thread until the PR merges or closes.
36
+ Retain a user-taken-over T3 thread and never send cleanup commands to it, including `t3_thread_organize` settle or archive.
37
+ Never remove the current driver, current pass, an active or unknown thread, an unsettled descendant, or a resource with ambiguous ownership.
38
+ At final settlement, no eligible non-driver thread or worktree remains.
39
+ At final settlement include run-owned worktrees in other repositories of the
40
+ same project and report each remaining resource as a hold with its reason. Recorded
41
+ projectId, threadId/runId, delegated taskId/childThreadId/childRunId, attempt key,
42
+ checkout path and revision decide ownership. Idle alone never proves exit.
43
+
44
+ Read back the durable receipt and evidence manifest before removing source copies or the worktree; missing or differing readback holds removal.
45
+ Evidence already outside the checkout needs no copy; read back its receipt and
46
+ supporting files. Keep settlement, evidence preservation, thread archival,
47
+ worktree removal and branch retirement as separate receipts.
48
+ Remove only the exact recorded path with `git worktree remove <path>` without force, then read back `git worktree list --porcelain` to verify absence.
49
+ Branches with a remote counterpart are never deleted.
50
+ Use exact `git branch -d <branch>` only for a proven run-owned local-only branch whose tip is reachable from the preserved candidate, verified salvage ref, or confirmed remote PR head; unique or unknown commits hold branch deletion.
51
+ Never delete the author branch for reviewer cleanup. Verify ref absence; a failed
52
+ or uncertain operation preserves the resource and resumes from its last receipt.
53
+
54
+ ## Owned schedule retirement
55
+
56
+ At Close-out reconcile recorded scheduledTaskIds with `list_scheduled_tasks`:
57
+ the owning run must have created each exact per-run watch selected for retirement.
58
+ Use `delete_scheduled_task` by exact ID and verify absence via
59
+ `list_scheduled_tasks` once nothing remains unsettled. An uncertain result holds
60
+ and retains the recorded ID. Cross-run retirement of an owned per-run watch
61
+ requires that its run is closed or every watched PR is merged or closed.
62
+ Read-only observers report to the driver; they do not remove schedules or worktrees.
63
+ Never remove the durable review-manager schedule, its lane resources, or a
64
+ user-created schedule through ordinary run settlement. See
65
+ [Review manager](automations.md) for scheduled-task health and lane policy.
66
+ Schedule deletion is distinct from thread archival and Git worktree removal.
71
67
 
72
68
  ## Preserve before removal
73
69
 
74
- Workers write reports, probes, logs, evidence, and scratch to the private
75
- `<run dir>/evidence/<dispatch>/` directory outside the disposable worktree.
76
- Name the exact directory in the dispatch brief and completion receipt. Peer
77
- reviewers use separate worktrees and separate evidence folders; neither reads
78
- the other's first-pass work. Authors commit the candidate before reporting
79
- done. Workers never push; the driver publishes under candidate-publication.
80
-
81
- If a completed eligible worktree has uncommitted or unpublished content,
82
- salvage before removal: create a salvage ref in that worktree, run `git add -A`
83
- and commit everything on that ref, write a `git bundle` for it into the private
84
- run folder, then run `git bundle verify`. Record the bundle path, bundle SHA-256, and salvage commit SHA in the
85
- receipt before removing the worktree. A failed verify holds the worktree. Never
86
- salvage an author worktree before merge or closure. Ignored non-cache files
87
- (anything other than known build and dependency caches), submodule changes, and content outside
88
- the worktree hold instead of being salvaged. Preserve any ambiguous source or
89
- publication state. Recheck the native owner and liveness immediately before
90
- removal, and use exact native worktree removal without force.
70
+ Workers write reports, probes, logs, evidence and scratch to private
71
+ `<run>/evidence/<key>/` outside disposable worktrees. Name the exact directory
72
+ in the brief and completion receipt. Peer reviewers use separate detached
73
+ checkouts and separate evidence folders with no first-pass cross-read. Authors commit
74
+ before reporting done; workers never push. The driver publishes under
75
+ [candidate publication](candidate-publication.md).
76
+
77
+ Before removal of an eligible non-author worktree with uncommitted or unpublished content, create a salvage ref, run `git add -A`, commit on that ref, write a `git bundle` into the private run folder, and run `git bundle verify`.
78
+ Ignored non-cache files must be explicitly classified and preserved in a verified salvage bundle before removal; unknown files hold the worktree.
79
+ Stage each classified ignored file separately with `git add -f -- <exact path>`
80
+ on the salvage ref before the commit; known build and dependency caches are
81
+ excluded. Verify the bundle includes every classified file's bytes and salvage
82
+ commit, beyond the bundle's structural verification.
83
+ Record the bundle path, bundle SHA-256, and salvage commit SHA in the receipt before removing the worktree.
84
+ A failed bundle verify holds the worktree.
85
+ Never salvage an author worktree before merge or closure.
86
+ Submodule changes and content outside the worktree hold instead of being salvaged.
87
+ Preserve ambiguous source or publication state. Recheck native owner, terminal
88
+ run evidence and no-writer proof immediately before archival and removal.
89
+ A verified bundle changes preservation classification, not ownership or liveness.
91
90
 
92
91
  ## Driver-start orphan sweep
93
92
 
94
- Drivers sweep on phase-skill entry. A scheduled pass that owns its lane with
95
- recorded cleanup authority runs the driver-start orphan sweep after predecessor
96
- terminal cleanup, scoped to repositories listed in its run record plus
97
- registered repositories on this host containing eligible settled resources of
98
- another Axstack run. On phase-skill entry, scope the sweep to the current repository
99
- and the per-run worktrees in other repositories recorded in the driver's
100
- run records, and registered repositories on this host containing eligible settled
101
- resources of another Axstack run. Inspect other Axstack run records on this host
102
- too; a live owning run does not protect its settled
103
- reviewer worktree or merged author after preservation checks. If Orca is unreachable,
104
- report one line and continue the phase; an unreachable host holds only its
105
- items. A sweep may remove resources of ANY Axstack run on this host after salvage
106
- when every owning Dispatch and descendant is settled (completed or failed), its
107
- release is confirmed or `release_unknown`, no agent is working, the exact terminal
108
- has had no output for at least 60 minutes, ownership and liveness are rechecked
109
- from a fresh native list, and evidence is durable. Remove descendants first.
110
- For a settled worker dispatched into a shared or driver worktree, close its
111
- exact terminal individually under the same 60-minute quiet rule: native close,
112
- else the guarded `/quit` + `exit` fallback above. Never close the live coordinator
113
- session's terminal for any run (the driver's own terminal) or a user chat.
114
- Align/Spec adviser sessions reused between rounds remain until
115
- their owning phase approves or stops; the quiet rule still applies then.
116
- For a merged or closed PR author with an unreleased settled Dispatch, request
117
- native worker release first, record its result, then apply the sweep guards.
118
- An author worktree of a merged or closed PR is sweep-eligible when its head
119
- commit is retrievable from the forge (for example, the PR's recorded head or a
120
- remote branch contains it); unverifiable state is a hold. For an eligible
121
- author worktree, salvage first if dirty, under the preservation guards above.
122
- Phase-skill entry drivers report sweep results and holds in chat and run record.
123
- A scheduled review-manager pass records sweep results and holds in its
124
- continuity record's Open holds table; a cleanup-authorized watch pass records
125
- them in its own continuity Open holds table. Both are silent when nothing was removed.
126
- Branches with a remote counterpart are never deleted. List live or unsettled
127
- work, genuine `user_takeover`, and items without provable Axstack provenance in
128
- one table with their reason; do not remove them.
129
- Dirty or unpublished work follows the salvage and publication guards above;
130
- open-PR authors remain protected until merge or close. Record each removed and
131
- held resource in the pass continuity (scheduled) or run record (driver start).
132
-
133
- ## Native bookkeeping exceptions
134
-
135
- A repair into an existing author terminal may be labelled `user_takeover` or
136
- `external_terminal`. Treat a repair-labelled terminal as run-owned only when
137
- the run record contains the exact repair Dispatch ID for that terminal;
138
- otherwise hold. Report the mislabel to Orca upstream through the driver.
139
- For `release_unknown`, reconcile with native worker inspection. If a fresh
140
- native terminal list confirms the terminal is gone, record the readback and
141
- proceed; otherwise hold unless inspection proves a settled, quiet worker. Then
142
- close only that worker's exact terminal under the sweep rule and confirm its
143
- absence. Unknown liveness holds.
144
-
145
- ## Known Orca issues
146
-
147
- Mark these for upstream reporting: in Orca 1.4.209 desktop, `orca terminal
148
- close` on an agent terminal returns `runtime_error` with `Error invoking remote
149
- method 'session:set': TypeError: Cannot convert undefined or null to object`.
150
- Dispatch into an existing terminal marks the worker retained/`user_takeover`.
151
- Cross-repository `--parent-worktree` is silently dropped.
152
-
153
- ## Readable sidebar
154
-
155
- Set the Orca display name with `orca worktree set --display-name` when creating
156
- each run worktree. Use run first, then role, space-separated: `<run> driver`,
157
- `<run> author #<pr>`, and `<run> review #<pr> r<n>`. Peer reviewers append
158
- `primary` or `secondary`; authored reviewers have no `primary` or `secondary`
159
- qualifier, and authors carry no round number. Use the task ID in
160
- place of `#<pr>` before a PR number exists, then update the name when assigned.
161
-
162
- Set `--comment` at dispatch, at settlement with the verdict and short SHA, and
163
- on a hold with its reason. Use `--workspace-status` for the coarse state and
164
- the comment for detail; never set `--workspace-status completed` for a hold.
165
-
166
- Use worktree parentage only to present ownership where supported: reviewer under
167
- its author, author under its driver in the same repository. Dependency order
168
- lives in names and `gh stack`. Do not rely on cross-repository parents, which
169
- Orca silently dropped, or remote `new-child`, which is invalid. Native Run,
170
- Task, and Dispatch receipts remain the source of ownership truth.
93
+ Drivers sweep on phase-skill entry. The sweep inventory is fully paginated project-scoped T3 threads plus `git worktree list --porcelain` plus run records.
94
+ Use projectId and exact recorded identities, not title substrings or age, to
95
+ prove Axstack provenance; include per-run worktrees in other repositories only
96
+ when the same project's records identify them. User-created threads and other projects' items are reported, never touched.
97
+ Report an unreachable T3 host and continue the phase; hold only its items.
98
+ Incomplete inventory holds affected eligibility; it never establishes absence.
99
+
100
+ Cross-run sweep eligibility for another Axstack run's settled reviewer or merged
101
+ or closed author in this project requires all attempts and descendants settled,
102
+ every agent inactive, the exact thread quiet for at least 60 minutes,
103
+ and fresh native state proving ownership and liveness with durable evidence.
104
+ Remove descendants first. Settled reviewer eligibility is independent of whether
105
+ its owning run is live. For a merged or closed author, prove its head is
106
+ retrievable from the forge (recorded PR head or remote branch); unverifiable state holds. Salvage an
107
+ eligible dirty author first under the preservation guards above. Align/Spec
108
+ advisers reused between rounds remain until their owning phase approves or stops.
109
+ Never archive active or waiting workers, or sweep solely because they are idle.
110
+
111
+ Record removed and held resources in the private run record, with exact
112
+ identities and resume conditions; phase-entry drivers also report in chat.
113
+ Cleanup-authorized scheduled passes record their own sweep results in continuity
114
+ Open holds. They stay silent when nothing was removed. List live or unsettled
115
+ work, user-taken-over threads, unknown provenance, user-created threads and
116
+ other-project items in one table with reasons. Age never grants deletion authority.
@@ -0,0 +1,120 @@
1
+ #!/usr/bin/env bun
2
+ import { readFileSync } from 'node:fs';
3
+
4
+ const queryFields = `
5
+ number headRefOid baseRefOid body isDraft state mergeable
6
+ commits(last: 1) { pageInfo { hasNextPage } nodes { commit {
7
+ statusCheckRollup { contexts(first: 100) { pageInfo { hasNextPage } nodes {
8
+ ... on CheckRun { id name status conclusion startedAt completedAt
9
+ checkSuite { app { slug } workflowRun { databaseId runNumber runAttempt } } }
10
+ ... on StatusContext { id context state createdAt }
11
+ } } }
12
+ } } }
13
+ reviews(first: 100) { pageInfo { hasNextPage } nodes {
14
+ id state body updatedAt submittedAt author { login }
15
+ } }
16
+ reviewRequests(first: 100) { pageInfo { hasNextPage } nodes {
17
+ requestedReviewer { ... on User { login } ... on Team { slug } }
18
+ } }
19
+ comments(first: 100) { pageInfo { hasNextPage } nodes { id body updatedAt isMinimized } }
20
+ reviewThreads(first: 100) { pageInfo { hasNextPage } nodes {
21
+ id isResolved isCollapsed comments(first: 100) {
22
+ pageInfo { hasNextPage } nodes { id body updatedAt }
23
+ }
24
+ } }
25
+ labels(first: 100) { pageInfo { hasNextPage } nodes { id name } }
26
+ `;
27
+
28
+ function options(argv) {
29
+ const result = {};
30
+ for (let i = 0; i < argv.length; i += 2) {
31
+ if (!['--input', '--watermark', '--repo', '--prs'].includes(argv[i]) || !argv[i + 1] || result[argv[i]]) {
32
+ throw new Error(`invalid argument: ${argv[i] ?? '(end)'}`);
33
+ }
34
+ result[argv[i]] = argv[i + 1];
35
+ }
36
+ if (!result['--watermark'] || (!result['--input'] && (!result['--repo'] || !result['--prs']))) {
37
+ throw new Error('expected --watermark path and either --input path or --repo owner/name --prs 1,2');
38
+ }
39
+ return result;
40
+ }
41
+
42
+ function fetchCurrent(args) {
43
+ if (args['--input']) return JSON.parse(readFileSync(args['--input'], 'utf8'));
44
+ const match = /^([\w.-]+)\/([\w.-]+)$/.exec(args['--repo']);
45
+ const numbers = args['--prs'].split(',').map(Number);
46
+ if (!match || !numbers.length || numbers.some((number) => !Number.isSafeInteger(number) || number < 1)) {
47
+ throw new Error('expected valid --repo owner/name and --prs 1,2');
48
+ }
49
+ const selections = numbers.map((number, index) => `pr${index}: pullRequest(number: ${number}) { ${queryFields} }`);
50
+ const query = `query { repository(owner: ${JSON.stringify(match[1])}, name: ${JSON.stringify(match[2])}) { ${selections.join('\n')} } }`;
51
+ const result = Bun.spawnSync(['gh', 'api', 'graphql', '-f', `query=${query}`], { stdout: 'pipe', stderr: 'pipe' });
52
+ if (result.exitCode !== 0) throw new Error(`GitHub GraphQL failed: ${result.stderr.toString().trim()}`);
53
+ return JSON.parse(result.stdout.toString());
54
+ }
55
+
56
+ function nodes(connection, field) {
57
+ if (!connection || connection.pageInfo?.hasNextPage !== false || !Array.isArray(connection.nodes)) {
58
+ throw new Error(`incomplete ${field} connection`);
59
+ }
60
+ return connection.nodes;
61
+ }
62
+
63
+ const digest = (body) => new Bun.CryptoHasher('sha256').update(body ?? '').digest('hex');
64
+ const ordered = (items) => items.sort((a, b) => JSON.stringify(a).localeCompare(JSON.stringify(b)));
65
+ const bodyItem = ({ body, ...rest }) => ({ ...rest, bodyDigest: digest(body) });
66
+
67
+ function snapshot(response) {
68
+ if (response.errors?.length || !response.data?.repository) throw new Error('incomplete GraphQL response');
69
+ const prs = {};
70
+ for (const pr of Object.values(response.data.repository)) {
71
+ if (!pr || !Number.isSafeInteger(pr.number) || !pr.headRefOid || !pr.baseRefOid) {
72
+ throw new Error('incomplete pull request');
73
+ }
74
+ const commits = nodes(pr.commits, 'commits');
75
+ const rollup = commits[0]?.commit?.statusCheckRollup;
76
+ const checks = rollup ? nodes(rollup.contexts, 'checks') : [];
77
+ const threads = nodes(pr.reviewThreads, 'reviewThreads').map((thread) => ({
78
+ id: thread.id, isResolved: thread.isResolved, isCollapsed: thread.isCollapsed,
79
+ comments: ordered(nodes(thread.comments, 'thread comments').map(bodyItem)),
80
+ }));
81
+ prs[pr.number] = {
82
+ headRefOid: pr.headRefOid, baseRefOid: pr.baseRefOid,
83
+ bodyDigest: digest(pr.body), isDraft: pr.isDraft, state: pr.state, mergeable: pr.mergeable,
84
+ checks: ordered(checks),
85
+ reviews: ordered(nodes(pr.reviews, 'reviews').map(bodyItem)),
86
+ reviewRequests: ordered(nodes(pr.reviewRequests, 'reviewRequests')),
87
+ comments: ordered(nodes(pr.comments, 'comments').map(bodyItem)),
88
+ reviewThreads: ordered(threads),
89
+ labels: ordered(nodes(pr.labels, 'labels')),
90
+ };
91
+ }
92
+ return prs;
93
+ }
94
+
95
+ try {
96
+ const args = options(process.argv.slice(2));
97
+ const current = snapshot(fetchCurrent(args));
98
+ let previous = {};
99
+ try {
100
+ previous = JSON.parse(readFileSync(args['--watermark'], 'utf8'));
101
+ if (previous?.data) previous = snapshot(previous);
102
+ }
103
+ catch (error) { if (error.code !== 'ENOENT') throw error; }
104
+ if (!previous || typeof previous !== 'object' || Array.isArray(previous)) throw new Error('invalid watermark');
105
+ const changes = [];
106
+ for (const number of new Set([...Object.keys(previous), ...Object.keys(current)])) {
107
+ const before = previous[number] ?? {};
108
+ const after = current[number] ?? {};
109
+ const fields = Object.keys({ ...before, ...after }).filter((key) => JSON.stringify(before[key]) !== JSON.stringify(after[key]));
110
+ if (fields.length) changes.push({ number: Number(number), fields });
111
+ }
112
+ if (changes.length) {
113
+ // The driver saves this watermark only after it has reconciled the deltas and pending local work.
114
+ console.log(JSON.stringify({ changes, watermark: current }));
115
+ process.exitCode = 10;
116
+ }
117
+ } catch (error) {
118
+ console.error(`PR digest incomplete: ${error.message}`);
119
+ process.exitCode = 2;
120
+ }
@@ -0,0 +1,102 @@
1
+ import { readFileSync } from 'node:fs';
2
+
3
+ const bindings = { codex: 'codex', claude: 'claudeAgent', grok: 'grok', antigravity: 'antigravity' };
4
+ const effortIds = { codex: 'reasoningEffort', claude: 'effort', grok: 'reasoningEffort' };
5
+ const nonempty = (value) => typeof value === 'string' && value.trim().length > 0;
6
+
7
+ function argumentsFrom(args) {
8
+ const options = { excluded: [] };
9
+ for (let i = 0; i < args.length; i += 2) {
10
+ const flag = args[i];
11
+ if (!['--provider', '--capabilities', '--class', '--model', '--effort', '--exclude'].includes(flag)) {
12
+ throw new Error(`unexpected argument ${flag}`);
13
+ }
14
+ const value = args[i + 1];
15
+ if (!nonempty(value) || value.startsWith('--')) throw new Error(`expected value after ${flag}`);
16
+ if (flag === '--exclude') options.excluded.push(value);
17
+ else {
18
+ const key = flag.slice(2);
19
+ if (key in options) throw new Error(`duplicate argument ${flag}`);
20
+ options[key] = value;
21
+ }
22
+ }
23
+ if (!options.provider || !options.capabilities || !options.effort) {
24
+ throw new Error('expected --provider codex|claude|grok|antigravity --capabilities path --effort level');
25
+ }
26
+ return options;
27
+ }
28
+
29
+ function providerFrom(catalog, provider) {
30
+ if (!Array.isArray(catalog?.providers) || catalog.providers.some((p) =>
31
+ !p || !nonempty(p.providerInstanceId) || !nonempty(p.driverKind))) {
32
+ throw new Error('malformed capabilities providers');
33
+ }
34
+ const matches = catalog.providers.filter((p) => p.providerInstanceId === bindings[provider]);
35
+ if (matches.length !== 1 || matches[0].driverKind !== bindings[provider]) {
36
+ throw new Error(`unknown or ambiguous provider instance ${bindings[provider]}`);
37
+ }
38
+ const instance = matches[0];
39
+ if (!Array.isArray(instance.models) || instance.models.some((model) =>
40
+ !model || !nonempty(model.id) || !Array.isArray(model.options) || model.options.some((option) =>
41
+ !option || !nonempty(option.id) || (option.options !== undefined &&
42
+ (!Array.isArray(option.options) || option.options.some((value) => !value || !nonempty(value.id))))))) {
43
+ throw new Error('malformed capabilities models or options');
44
+ }
45
+ return instance;
46
+ }
47
+
48
+ function modelFrom(models, { provider, model: pin, class: modelClass, excluded }) {
49
+ if (pin && pin !== 'null') {
50
+ const model = models.find((model) => model.id === pin && !excluded.includes(model.id));
51
+ if (!model) throw new Error(`missing requested model ${pin}`);
52
+ return model;
53
+ }
54
+ if (!pin && !modelClass) throw new Error('expected --class or --model (use --model null for a null role)');
55
+ if (!modelClass) {
56
+ if (['codex', 'claude'].includes(provider)) {
57
+ throw new Error(`intentional absence for ${provider}: missing model and class`);
58
+ }
59
+ // Launch-by-agent-ID providers bind the first listed ID; exclusions never select a substitute.
60
+ const model = models[0];
61
+ if (!model || excluded.includes(model.id)) throw new Error('missing first listed model');
62
+ return model;
63
+ }
64
+ if (!['codex', 'claude'].includes(provider) || !/^[a-z]+$/.test(modelClass)) {
65
+ throw new Error(`unsupported model class ${modelClass} for ${provider}`);
66
+ }
67
+ const pattern = provider === 'codex'
68
+ ? new RegExp(`^gpt-(\\d+(?:\\.\\d+)*)-${modelClass}$`)
69
+ : new RegExp(`^claude-${modelClass}-(\\d+-\\d+)$`);
70
+ const candidates = models.filter((model) => pattern.test(model.id) && !excluded.includes(model.id));
71
+ candidates.sort((a, b) => {
72
+ const left = a.id.match(pattern)[1].split(/[.-]/).map(Number);
73
+ const right = b.id.match(pattern)[1].split(/[.-]/).map(Number);
74
+ for (let i = 0; i < Math.max(left.length, right.length); i++) {
75
+ const difference = (right[i] ?? 0) - (left[i] ?? 0);
76
+ if (difference) return difference;
77
+ }
78
+ return 0;
79
+ });
80
+ if (!candidates.length) throw new Error(`no matching ${modelClass} model for ${provider}`);
81
+ return candidates[0];
82
+ }
83
+
84
+ try {
85
+ const options = argumentsFrom(process.argv.slice(2));
86
+ const { provider, capabilities: path, effort } = options;
87
+ if (!Object.hasOwn(bindings, provider)) throw new Error(`unknown provider ${provider}`);
88
+ const catalog = JSON.parse(readFileSync(path, 'utf8'));
89
+ const instance = providerFrom(catalog, provider);
90
+ const model = modelFrom(instance.models, options);
91
+ const effortOptions = model.options.filter((option) => effortIds[provider]
92
+ ? option.id === effortIds[provider] : ['reasoningEffort', 'effort'].includes(option.id));
93
+ if (effortOptions.length !== 1 || !effortOptions[0].options?.some((value) => value.id === effort)
94
+ || (provider === 'grok' && effort === 'max')) {
95
+ throw new Error(`unsupported effort ${effort} for ${provider}/${model.id}`);
96
+ }
97
+ console.log(JSON.stringify({ provider, providerInstanceId: instance.providerInstanceId,
98
+ model: model.id, effortOption: { id: effortOptions[0].id, value: effort }, source: 'capabilities', path }));
99
+ } catch (error) {
100
+ console.error(`model capabilities resolution hold: ${error.message}`);
101
+ process.exitCode = 1;
102
+ }
@@ -5,6 +5,9 @@ description: When exploring or planning engineering work, use axstack-align to s
5
5
 
6
6
  # Align
7
7
 
8
+ For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
9
+ for phase continuation and holds.
10
+
8
11
  On driver entry, sweep under [Workspace hygiene](../axstack/references/workspace-hygiene.md); dispatched workers do not sweep.
9
12
  For every dispatch brief, name its private `<run dir>/evidence/<dispatch>/` folder.
10
13
 
@@ -34,7 +37,7 @@ substantial; apply routing's existing size reassessment rule.
34
37
  tools before asking the user. Separate facts from preferences, name evidence
35
38
  gaps, and map which decisions unlock others. When a fact needed for the
36
39
  frontier is not derivable from the local repo or docs by ordinary reading,
37
- dispatch `axstack-research` branches through Orca by source type:
40
+ dispatch `axstack-research` branches through T3 `delegate_task` by source type:
38
41
  requirements, code, web, and, once configured, X. Give one owner per branch,
39
42
  use cross-harness routes where the roles allow, and require a cited note per
40
43
  the research skill's source standards. The driver folds verified claims into
@@ -73,12 +76,17 @@ material disagreement remains, then surface the choices to the user. Never
73
76
  fabricate consensus or impersonate a role.
74
77
 
75
78
  Immediately before the first actual adviser dispatch, load and follow
76
- [Orca runtime](../axstack/references/orca-runtime.md). Reuse each adviser
79
+ [T3 runtime](../axstack/references/t3-runtime.md). Reuse each adviser
77
80
  session and settled receipt; consult only the changed frontier and reuse
78
81
  unchanged receipts. Record compact adviser evidence, the driver's assessment,
79
82
  and user-resolved choices for `axstack-spec`. If either adviser is unavailable,
80
83
  hold Align; safe fact work may continue without substitution.
81
84
 
85
+ An optional adviser note may be deferred or rejected in a `Decisions` row with
86
+ the draft unchanged; it needs no new adviser pair. Changed draft text, a
87
+ blocking finding, or a high-stakes decision requires fresh receipts on the new
88
+ revision.
89
+
82
90
  ## Arena for hard-to-reverse design choices
83
91
 
84
92
  Critique of one draft anchors every reader to that draft's shape. Rung 2 designs
@@ -122,10 +130,15 @@ an arena. Small or routine questions never enter the arena.
122
130
  Record the synthesis note (base, grafts and their source candidate, rejections,
123
131
  dropouts, judge verdicts per round) as `Decisions` rows in the
124
132
  [run record](../axstack/references/run-record.md). Load
125
- [Orca runtime](../axstack/references/orca-runtime.md) immediately before the
126
- first candidate or judge dispatch. If any configured candidate or judge seat
127
- required for that round is unavailable at launch or returns a failed receipt,
128
- hold that question without substitution, record the gap, and ask: the user decides whether to proceed without it.
133
+ [T3 runtime](../axstack/references/t3-runtime.md) immediately before the
134
+ first candidate or judge dispatch. If an optional Grok or Antigravity candidate
135
+ malfunctions (launch failure, trust/login prompt, or prompt block), fence it,
136
+ record `absent (<reason>)`, name it once in the next
137
+ read-back, and continue with available candidates without relay or substitution.
138
+ A required adviser, candidate, or judge unavailable at launch or returning a
139
+ failed receipt holds that question without substitution; record the gap and ask
140
+ whether to proceed. In mixed fan-out retain at least one Codex and one Claude
141
+ seat, or hold the affected question.
129
142
  For an uncertain dispatch, reconcile natively; it is never treated as absent.
130
143
  Unaffected fact work and questions continue.
131
144
 
@@ -166,7 +179,7 @@ record or spec. Read-only scope keeps proposed documentation in the permitted
166
179
  private record or response. Documentation is neither implementation nor spec
167
180
  approval; record chosen document names and paths once per run.
168
181
 
169
- ## Read back, classify, and stop
182
+ ## Read back, classify, and route
170
183
 
171
184
  1. Read back the decisions, constraints, exclusions, and remaining evidence
172
185
  gaps. For substantial work, this summary becomes part of the draft spec in
@@ -186,9 +199,10 @@ approval; record chosen document names and paths once per run.
186
199
  documentation pointers without adding another runtime. Return the compact
187
200
  scope and record pointer in the current chat. Native transfer is separate:
188
201
  use it only when the user explicitly requests transfer, loading
189
- [Orca runtime](../axstack/references/orca-runtime.md) immediately before
202
+ [T3 runtime](../axstack/references/t3-runtime.md) immediately before
190
203
  actual dispatch. Alignment completion never dispatches a recipient.
191
204
 
192
- Alignment stops for both sizes only when the handoff is usable, its next scope
193
- identity is explicit, and execution has not started. The user invokes
194
- `axstack-implement` to execute.
205
+ Alignment completes for both sizes only when the handoff is usable and its next
206
+ scope identity is explicit. An eligible delivery run continues under Autopilot;
207
+ an explicit stop-after-Align request ends here. Substantial work continues to
208
+ Spec, and small work continues from its small-change intent to Implement.