axstack 0.20.30 → 0.21.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +24 -23
- package/bin/axstack.js +18 -5
- package/docs/installation.md +101 -46
- package/docs/workflows.md +179 -117
- package/package.json +3 -3
- package/profiles/presets/claude-only.json +46 -46
- package/profiles/presets/codex-only.json +50 -50
- package/profiles/presets/mixed.json +59 -59
- package/skills/axstack/references/automations.md +127 -137
- package/skills/axstack/references/autopilot.md +121 -0
- package/skills/axstack/references/candidate-publication.md +13 -8
- package/skills/axstack/references/contracts.md +10 -4
- package/skills/axstack/references/diligence.md +3 -1
- package/skills/axstack/references/evidence-archive.md +38 -33
- package/skills/axstack/references/lifecycle.md +64 -50
- package/skills/axstack/references/review-manager-prompt.md +13 -11
- package/skills/axstack/references/role-roster.md +19 -9
- package/skills/axstack/references/routing.md +33 -18
- package/skills/axstack/references/run-record.md +36 -15
- package/skills/axstack/references/t3-runtime.md +234 -0
- package/skills/axstack/references/test-audit-weekly.md +62 -0
- package/skills/axstack/references/test-value.md +120 -0
- package/skills/axstack/references/ui-verification.md +5 -1
- package/skills/axstack/references/workspace-hygiene.md +102 -156
- package/skills/axstack/scripts/pr-digest.js +120 -0
- package/skills/axstack/scripts/resolve-models.js +102 -0
- package/skills/axstack-align/SKILL.md +25 -11
- package/skills/axstack-audit/SKILL.md +22 -5
- package/skills/axstack-audit/references/record.md +1 -1
- package/skills/axstack-cleanup/SKILL.md +69 -87
- package/skills/axstack-debug/SKILL.md +1 -1
- package/skills/axstack-explain/SKILL.md +1 -1
- package/skills/axstack-explain/references/visual-qa.md +2 -0
- package/skills/axstack-implement/SKILL.md +76 -26
- package/skills/axstack-improve/SKILL.md +24 -4
- package/skills/axstack-relay/SKILL.md +16 -7
- package/skills/axstack-research/SKILL.md +11 -4
- package/skills/axstack-review/SKILL.md +42 -32
- package/skills/axstack-spec/SKILL.md +23 -14
- package/skills/axstack-tickets/SKILL.md +13 -11
- package/skills/axstack-watch/SKILL.md +117 -34
- package/skills/axstack-watch/references/watch-runtime.md +61 -69
- package/src/capabilities.js +33 -69
- package/src/installer.js +9 -1
- package/src/instructions.js +9 -4
- package/src/roles.js +38 -10
- package/skills/axstack/references/orca-runtime.md +0 -183
- package/skills/axstack/scripts/trust-path.js +0 -123
|
@@ -1,170 +1,116 @@
|
|
|
1
|
-
# Workspace hygiene for driver-owned
|
|
1
|
+
# Workspace hygiene for driver-owned T3 runs
|
|
2
2
|
|
|
3
|
-
This is a prompt contract
|
|
4
|
-
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
holds only the affected resource.
|
|
3
|
+
This is a prompt contract, not a cleanup daemon. Follow [T3 runtime](t3-runtime.md)
|
|
4
|
+
for native identities, terminal run evidence and schema. Dispatched workers never
|
|
5
|
+
sweep or remove another session. A scheduled pass with recorded cleanup authority
|
|
6
|
+
acts as its lane's driver; read-only observers only report leftovers. Uncertain
|
|
7
|
+
ownership, liveness or evidence holds only the affected resource.
|
|
8
|
+
Record each decision and native readback in the private run record.
|
|
10
9
|
|
|
11
10
|
## Safe deletion
|
|
12
11
|
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
12
|
+
Before use, commands must scope `TMPDIR` to an owned 0700 directory under the system temp directory, never under `$HOME`, named from the dispatch key and recorded in the receipt.
|
|
13
|
+
Validate its real path, absence of symlinks and ownership before use and cleanup; remove it afterwards by literal absolute path.
|
|
14
|
+
Evidence files still go to the private `<run>/evidence/<key>/` folder.
|
|
15
|
+
|
|
16
|
+
Every shell deletion targets a literal absolute path or a `${VAR:?}`-guarded expansion, only inside the worker's own evidence folder, `TMPDIR`, or worktree.
|
|
17
|
+
For example, `rm -rf -- "${EV:?}/mut"` requires a validated owned evidence path.
|
|
18
|
+
Never use a bare `$VAR`, a glob on a variable, `/`, `HOME`, or a shared root as a deletion target.
|
|
19
|
+
Prefer `git clean -- <exact prefix>` or tool-native cleanup. A safety prompt that
|
|
18
20
|
still appears is a hold; agents do not answer it.
|
|
19
21
|
|
|
22
|
+
## Project preflight
|
|
23
|
+
|
|
24
|
+
`worktreeCleanup` must be `off` for every Axstack project.
|
|
25
|
+
Preflight reads back `worktreeCleanup` via `t3_project_read` where exposed; otherwise record a limitation pointing to the documented installation setup step.
|
|
26
|
+
Automatic deletion cannot replace evidence readback or salvage. Do not change
|
|
27
|
+
project settings without recorded host-mutation authority.
|
|
28
|
+
|
|
20
29
|
## Settlement
|
|
21
30
|
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
At final settlement, no eligible non-driver
|
|
30
|
-
|
|
31
|
-
resource as a hold with its reason. Recorded
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
`orca automations remove <id>` on its exact ID, and verify absence by native
|
|
59
|
-
readback. The observer only disables and reports; the driver removes its own
|
|
60
|
-
run's automation. Only a retired owned per-run watch's workspace may be removed
|
|
61
|
-
under these guards. Never remove the durable review manager and its dedicated
|
|
62
|
-
workspace or a user-created automation and its dedicated workspace.
|
|
63
|
-
|
|
64
|
-
Remove that dedicated workspace only after confirming its exact ownership, no
|
|
65
|
-
live terminal, a clean worktree, and a head on the remote. Recheck ownership and
|
|
66
|
-
liveness immediately before workspace removal. If dirty or unpublished, use the
|
|
67
|
-
salvage and bundle verification below before removal; failed verification or
|
|
68
|
-
uncertain liveness is a hold. The current scheduled pass cannot remove its own
|
|
69
|
-
workspace while its terminal is live; the driver or a later sweep finishes that
|
|
70
|
-
step. Record separate automation and workspace receipts.
|
|
31
|
+
Cleanup order is: settled descendants -> evidence readback -> salvage dirty or ignored non-cache content -> t3_thread_organize archive -> exact git worktree remove without force -> git branch -d for local-only branches.
|
|
32
|
+
Require terminal run evidence before `t3_thread_organize` settle or archive.
|
|
33
|
+
These metadata actions do not remove worktrees. Accept matching delegated task
|
|
34
|
+
or launched run completion under the runtime contract before settlement.
|
|
35
|
+
Keep the author worktree and thread until the PR merges or closes.
|
|
36
|
+
Retain a user-taken-over T3 thread and never send cleanup commands to it, including `t3_thread_organize` settle or archive.
|
|
37
|
+
Never remove the current driver, current pass, an active or unknown thread, an unsettled descendant, or a resource with ambiguous ownership.
|
|
38
|
+
At final settlement, no eligible non-driver thread or worktree remains.
|
|
39
|
+
At final settlement include run-owned worktrees in other repositories of the
|
|
40
|
+
same project and report each remaining resource as a hold with its reason. Recorded
|
|
41
|
+
projectId, threadId/runId, delegated taskId/childThreadId/childRunId, attempt key,
|
|
42
|
+
checkout path and revision decide ownership. Idle alone never proves exit.
|
|
43
|
+
|
|
44
|
+
Read back the durable receipt and evidence manifest before removing source copies or the worktree; missing or differing readback holds removal.
|
|
45
|
+
Evidence already outside the checkout needs no copy; read back its receipt and
|
|
46
|
+
supporting files. Keep settlement, evidence preservation, thread archival,
|
|
47
|
+
worktree removal and branch retirement as separate receipts.
|
|
48
|
+
Remove only the exact recorded path with `git worktree remove <path>` without force, then read back `git worktree list --porcelain` to verify absence.
|
|
49
|
+
Branches with a remote counterpart are never deleted.
|
|
50
|
+
Use exact `git branch -d <branch>` only for a proven run-owned local-only branch whose tip is reachable from the preserved candidate, verified salvage ref, or confirmed remote PR head; unique or unknown commits hold branch deletion.
|
|
51
|
+
Never delete the author branch for reviewer cleanup. Verify ref absence; a failed
|
|
52
|
+
or uncertain operation preserves the resource and resumes from its last receipt.
|
|
53
|
+
|
|
54
|
+
## Owned schedule retirement
|
|
55
|
+
|
|
56
|
+
At Close-out reconcile recorded scheduledTaskIds with `list_scheduled_tasks`:
|
|
57
|
+
the owning run must have created each exact per-run watch selected for retirement.
|
|
58
|
+
Use `delete_scheduled_task` by exact ID and verify absence via
|
|
59
|
+
`list_scheduled_tasks` once nothing remains unsettled. An uncertain result holds
|
|
60
|
+
and retains the recorded ID. Cross-run retirement of an owned per-run watch
|
|
61
|
+
requires that its run is closed or every watched PR is merged or closed.
|
|
62
|
+
Read-only observers report to the driver; they do not remove schedules or worktrees.
|
|
63
|
+
Never remove the durable review-manager schedule, its lane resources, or a
|
|
64
|
+
user-created schedule through ordinary run settlement. See
|
|
65
|
+
[Review manager](automations.md) for scheduled-task health and lane policy.
|
|
66
|
+
Schedule deletion is distinct from thread archival and Git worktree removal.
|
|
71
67
|
|
|
72
68
|
## Preserve before removal
|
|
73
69
|
|
|
74
|
-
Workers write reports, probes, logs, evidence
|
|
75
|
-
`<run
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
70
|
+
Workers write reports, probes, logs, evidence and scratch to private
|
|
71
|
+
`<run>/evidence/<key>/` outside disposable worktrees. Name the exact directory
|
|
72
|
+
in the brief and completion receipt. Peer reviewers use separate detached
|
|
73
|
+
checkouts and separate evidence folders with no first-pass cross-read. Authors commit
|
|
74
|
+
before reporting done; workers never push. The driver publishes under
|
|
75
|
+
[candidate publication](candidate-publication.md).
|
|
76
|
+
|
|
77
|
+
Before removal of an eligible non-author worktree with uncommitted or unpublished content, create a salvage ref, run `git add -A`, commit on that ref, write a `git bundle` into the private run folder, and run `git bundle verify`.
|
|
78
|
+
Ignored non-cache files must be explicitly classified and preserved in a verified salvage bundle before removal; unknown files hold the worktree.
|
|
79
|
+
Stage each classified ignored file separately with `git add -f -- <exact path>`
|
|
80
|
+
on the salvage ref before the commit; known build and dependency caches are
|
|
81
|
+
excluded. Verify the bundle includes every classified file's bytes and salvage
|
|
82
|
+
commit, beyond the bundle's structural verification.
|
|
83
|
+
Record the bundle path, bundle SHA-256, and salvage commit SHA in the receipt before removing the worktree.
|
|
84
|
+
A failed bundle verify holds the worktree.
|
|
85
|
+
Never salvage an author worktree before merge or closure.
|
|
86
|
+
Submodule changes and content outside the worktree hold instead of being salvaged.
|
|
87
|
+
Preserve ambiguous source or publication state. Recheck native owner, terminal
|
|
88
|
+
run evidence and no-writer proof immediately before archival and removal.
|
|
89
|
+
A verified bundle changes preservation classification, not ownership or liveness.
|
|
91
90
|
|
|
92
91
|
## Driver-start orphan sweep
|
|
93
92
|
|
|
94
|
-
Drivers sweep on phase-skill entry.
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
An author worktree of a merged or closed PR is sweep-eligible when its head
|
|
119
|
-
commit is retrievable from the forge (for example, the PR's recorded head or a
|
|
120
|
-
remote branch contains it); unverifiable state is a hold. For an eligible
|
|
121
|
-
author worktree, salvage first if dirty, under the preservation guards above.
|
|
122
|
-
Phase-skill entry drivers report sweep results and holds in chat and run record.
|
|
123
|
-
A scheduled review-manager pass records sweep results and holds in its
|
|
124
|
-
continuity record's Open holds table; a cleanup-authorized watch pass records
|
|
125
|
-
them in its own continuity Open holds table. Both are silent when nothing was removed.
|
|
126
|
-
Branches with a remote counterpart are never deleted. List live or unsettled
|
|
127
|
-
work, genuine `user_takeover`, and items without provable Axstack provenance in
|
|
128
|
-
one table with their reason; do not remove them.
|
|
129
|
-
Dirty or unpublished work follows the salvage and publication guards above;
|
|
130
|
-
open-PR authors remain protected until merge or close. Record each removed and
|
|
131
|
-
held resource in the pass continuity (scheduled) or run record (driver start).
|
|
132
|
-
|
|
133
|
-
## Native bookkeeping exceptions
|
|
134
|
-
|
|
135
|
-
A repair into an existing author terminal may be labelled `user_takeover` or
|
|
136
|
-
`external_terminal`. Treat a repair-labelled terminal as run-owned only when
|
|
137
|
-
the run record contains the exact repair Dispatch ID for that terminal;
|
|
138
|
-
otherwise hold. Report the mislabel to Orca upstream through the driver.
|
|
139
|
-
For `release_unknown`, reconcile with native worker inspection. If a fresh
|
|
140
|
-
native terminal list confirms the terminal is gone, record the readback and
|
|
141
|
-
proceed; otherwise hold unless inspection proves a settled, quiet worker. Then
|
|
142
|
-
close only that worker's exact terminal under the sweep rule and confirm its
|
|
143
|
-
absence. Unknown liveness holds.
|
|
144
|
-
|
|
145
|
-
## Known Orca issues
|
|
146
|
-
|
|
147
|
-
Mark these for upstream reporting: in Orca 1.4.209 desktop, `orca terminal
|
|
148
|
-
close` on an agent terminal returns `runtime_error` with `Error invoking remote
|
|
149
|
-
method 'session:set': TypeError: Cannot convert undefined or null to object`.
|
|
150
|
-
Dispatch into an existing terminal marks the worker retained/`user_takeover`.
|
|
151
|
-
Cross-repository `--parent-worktree` is silently dropped.
|
|
152
|
-
|
|
153
|
-
## Readable sidebar
|
|
154
|
-
|
|
155
|
-
Set the Orca display name with `orca worktree set --display-name` when creating
|
|
156
|
-
each run worktree. Use run first, then role, space-separated: `<run> driver`,
|
|
157
|
-
`<run> author #<pr>`, and `<run> review #<pr> r<n>`. Peer reviewers append
|
|
158
|
-
`primary` or `secondary`; authored reviewers have no `primary` or `secondary`
|
|
159
|
-
qualifier, and authors carry no round number. Use the task ID in
|
|
160
|
-
place of `#<pr>` before a PR number exists, then update the name when assigned.
|
|
161
|
-
|
|
162
|
-
Set `--comment` at dispatch, at settlement with the verdict and short SHA, and
|
|
163
|
-
on a hold with its reason. Use `--workspace-status` for the coarse state and
|
|
164
|
-
the comment for detail; never set `--workspace-status completed` for a hold.
|
|
165
|
-
|
|
166
|
-
Use worktree parentage only to present ownership where supported: reviewer under
|
|
167
|
-
its author, author under its driver in the same repository. Dependency order
|
|
168
|
-
lives in names and `gh stack`. Do not rely on cross-repository parents, which
|
|
169
|
-
Orca silently dropped, or remote `new-child`, which is invalid. Native Run,
|
|
170
|
-
Task, and Dispatch receipts remain the source of ownership truth.
|
|
93
|
+
Drivers sweep on phase-skill entry. The sweep inventory is fully paginated project-scoped T3 threads plus `git worktree list --porcelain` plus run records.
|
|
94
|
+
Use projectId and exact recorded identities, not title substrings or age, to
|
|
95
|
+
prove Axstack provenance; include per-run worktrees in other repositories only
|
|
96
|
+
when the same project's records identify them. User-created threads and other projects' items are reported, never touched.
|
|
97
|
+
Report an unreachable T3 host and continue the phase; hold only its items.
|
|
98
|
+
Incomplete inventory holds affected eligibility; it never establishes absence.
|
|
99
|
+
|
|
100
|
+
Cross-run sweep eligibility for another Axstack run's settled reviewer or merged
|
|
101
|
+
or closed author in this project requires all attempts and descendants settled,
|
|
102
|
+
every agent inactive, the exact thread quiet for at least 60 minutes,
|
|
103
|
+
and fresh native state proving ownership and liveness with durable evidence.
|
|
104
|
+
Remove descendants first. Settled reviewer eligibility is independent of whether
|
|
105
|
+
its owning run is live. For a merged or closed author, prove its head is
|
|
106
|
+
retrievable from the forge (recorded PR head or remote branch); unverifiable state holds. Salvage an
|
|
107
|
+
eligible dirty author first under the preservation guards above. Align/Spec
|
|
108
|
+
advisers reused between rounds remain until their owning phase approves or stops.
|
|
109
|
+
Never archive active or waiting workers, or sweep solely because they are idle.
|
|
110
|
+
|
|
111
|
+
Record removed and held resources in the private run record, with exact
|
|
112
|
+
identities and resume conditions; phase-entry drivers also report in chat.
|
|
113
|
+
Cleanup-authorized scheduled passes record their own sweep results in continuity
|
|
114
|
+
Open holds. They stay silent when nothing was removed. List live or unsettled
|
|
115
|
+
work, user-taken-over threads, unknown provenance, user-created threads and
|
|
116
|
+
other-project items in one table with reasons. Age never grants deletion authority.
|
|
@@ -0,0 +1,120 @@
|
|
|
1
|
+
#!/usr/bin/env bun
|
|
2
|
+
import { readFileSync } from 'node:fs';
|
|
3
|
+
|
|
4
|
+
const queryFields = `
|
|
5
|
+
number headRefOid baseRefOid body isDraft state mergeable
|
|
6
|
+
commits(last: 1) { pageInfo { hasNextPage } nodes { commit {
|
|
7
|
+
statusCheckRollup { contexts(first: 100) { pageInfo { hasNextPage } nodes {
|
|
8
|
+
... on CheckRun { id name status conclusion startedAt completedAt
|
|
9
|
+
checkSuite { app { slug } workflowRun { databaseId runNumber runAttempt } } }
|
|
10
|
+
... on StatusContext { id context state createdAt }
|
|
11
|
+
} } }
|
|
12
|
+
} } }
|
|
13
|
+
reviews(first: 100) { pageInfo { hasNextPage } nodes {
|
|
14
|
+
id state body updatedAt submittedAt author { login }
|
|
15
|
+
} }
|
|
16
|
+
reviewRequests(first: 100) { pageInfo { hasNextPage } nodes {
|
|
17
|
+
requestedReviewer { ... on User { login } ... on Team { slug } }
|
|
18
|
+
} }
|
|
19
|
+
comments(first: 100) { pageInfo { hasNextPage } nodes { id body updatedAt isMinimized } }
|
|
20
|
+
reviewThreads(first: 100) { pageInfo { hasNextPage } nodes {
|
|
21
|
+
id isResolved isCollapsed comments(first: 100) {
|
|
22
|
+
pageInfo { hasNextPage } nodes { id body updatedAt }
|
|
23
|
+
}
|
|
24
|
+
} }
|
|
25
|
+
labels(first: 100) { pageInfo { hasNextPage } nodes { id name } }
|
|
26
|
+
`;
|
|
27
|
+
|
|
28
|
+
function options(argv) {
|
|
29
|
+
const result = {};
|
|
30
|
+
for (let i = 0; i < argv.length; i += 2) {
|
|
31
|
+
if (!['--input', '--watermark', '--repo', '--prs'].includes(argv[i]) || !argv[i + 1] || result[argv[i]]) {
|
|
32
|
+
throw new Error(`invalid argument: ${argv[i] ?? '(end)'}`);
|
|
33
|
+
}
|
|
34
|
+
result[argv[i]] = argv[i + 1];
|
|
35
|
+
}
|
|
36
|
+
if (!result['--watermark'] || (!result['--input'] && (!result['--repo'] || !result['--prs']))) {
|
|
37
|
+
throw new Error('expected --watermark path and either --input path or --repo owner/name --prs 1,2');
|
|
38
|
+
}
|
|
39
|
+
return result;
|
|
40
|
+
}
|
|
41
|
+
|
|
42
|
+
function fetchCurrent(args) {
|
|
43
|
+
if (args['--input']) return JSON.parse(readFileSync(args['--input'], 'utf8'));
|
|
44
|
+
const match = /^([\w.-]+)\/([\w.-]+)$/.exec(args['--repo']);
|
|
45
|
+
const numbers = args['--prs'].split(',').map(Number);
|
|
46
|
+
if (!match || !numbers.length || numbers.some((number) => !Number.isSafeInteger(number) || number < 1)) {
|
|
47
|
+
throw new Error('expected valid --repo owner/name and --prs 1,2');
|
|
48
|
+
}
|
|
49
|
+
const selections = numbers.map((number, index) => `pr${index}: pullRequest(number: ${number}) { ${queryFields} }`);
|
|
50
|
+
const query = `query { repository(owner: ${JSON.stringify(match[1])}, name: ${JSON.stringify(match[2])}) { ${selections.join('\n')} } }`;
|
|
51
|
+
const result = Bun.spawnSync(['gh', 'api', 'graphql', '-f', `query=${query}`], { stdout: 'pipe', stderr: 'pipe' });
|
|
52
|
+
if (result.exitCode !== 0) throw new Error(`GitHub GraphQL failed: ${result.stderr.toString().trim()}`);
|
|
53
|
+
return JSON.parse(result.stdout.toString());
|
|
54
|
+
}
|
|
55
|
+
|
|
56
|
+
function nodes(connection, field) {
|
|
57
|
+
if (!connection || connection.pageInfo?.hasNextPage !== false || !Array.isArray(connection.nodes)) {
|
|
58
|
+
throw new Error(`incomplete ${field} connection`);
|
|
59
|
+
}
|
|
60
|
+
return connection.nodes;
|
|
61
|
+
}
|
|
62
|
+
|
|
63
|
+
const digest = (body) => new Bun.CryptoHasher('sha256').update(body ?? '').digest('hex');
|
|
64
|
+
const ordered = (items) => items.sort((a, b) => JSON.stringify(a).localeCompare(JSON.stringify(b)));
|
|
65
|
+
const bodyItem = ({ body, ...rest }) => ({ ...rest, bodyDigest: digest(body) });
|
|
66
|
+
|
|
67
|
+
function snapshot(response) {
|
|
68
|
+
if (response.errors?.length || !response.data?.repository) throw new Error('incomplete GraphQL response');
|
|
69
|
+
const prs = {};
|
|
70
|
+
for (const pr of Object.values(response.data.repository)) {
|
|
71
|
+
if (!pr || !Number.isSafeInteger(pr.number) || !pr.headRefOid || !pr.baseRefOid) {
|
|
72
|
+
throw new Error('incomplete pull request');
|
|
73
|
+
}
|
|
74
|
+
const commits = nodes(pr.commits, 'commits');
|
|
75
|
+
const rollup = commits[0]?.commit?.statusCheckRollup;
|
|
76
|
+
const checks = rollup ? nodes(rollup.contexts, 'checks') : [];
|
|
77
|
+
const threads = nodes(pr.reviewThreads, 'reviewThreads').map((thread) => ({
|
|
78
|
+
id: thread.id, isResolved: thread.isResolved, isCollapsed: thread.isCollapsed,
|
|
79
|
+
comments: ordered(nodes(thread.comments, 'thread comments').map(bodyItem)),
|
|
80
|
+
}));
|
|
81
|
+
prs[pr.number] = {
|
|
82
|
+
headRefOid: pr.headRefOid, baseRefOid: pr.baseRefOid,
|
|
83
|
+
bodyDigest: digest(pr.body), isDraft: pr.isDraft, state: pr.state, mergeable: pr.mergeable,
|
|
84
|
+
checks: ordered(checks),
|
|
85
|
+
reviews: ordered(nodes(pr.reviews, 'reviews').map(bodyItem)),
|
|
86
|
+
reviewRequests: ordered(nodes(pr.reviewRequests, 'reviewRequests')),
|
|
87
|
+
comments: ordered(nodes(pr.comments, 'comments').map(bodyItem)),
|
|
88
|
+
reviewThreads: ordered(threads),
|
|
89
|
+
labels: ordered(nodes(pr.labels, 'labels')),
|
|
90
|
+
};
|
|
91
|
+
}
|
|
92
|
+
return prs;
|
|
93
|
+
}
|
|
94
|
+
|
|
95
|
+
try {
|
|
96
|
+
const args = options(process.argv.slice(2));
|
|
97
|
+
const current = snapshot(fetchCurrent(args));
|
|
98
|
+
let previous = {};
|
|
99
|
+
try {
|
|
100
|
+
previous = JSON.parse(readFileSync(args['--watermark'], 'utf8'));
|
|
101
|
+
if (previous?.data) previous = snapshot(previous);
|
|
102
|
+
}
|
|
103
|
+
catch (error) { if (error.code !== 'ENOENT') throw error; }
|
|
104
|
+
if (!previous || typeof previous !== 'object' || Array.isArray(previous)) throw new Error('invalid watermark');
|
|
105
|
+
const changes = [];
|
|
106
|
+
for (const number of new Set([...Object.keys(previous), ...Object.keys(current)])) {
|
|
107
|
+
const before = previous[number] ?? {};
|
|
108
|
+
const after = current[number] ?? {};
|
|
109
|
+
const fields = Object.keys({ ...before, ...after }).filter((key) => JSON.stringify(before[key]) !== JSON.stringify(after[key]));
|
|
110
|
+
if (fields.length) changes.push({ number: Number(number), fields });
|
|
111
|
+
}
|
|
112
|
+
if (changes.length) {
|
|
113
|
+
// The driver saves this watermark only after it has reconciled the deltas and pending local work.
|
|
114
|
+
console.log(JSON.stringify({ changes, watermark: current }));
|
|
115
|
+
process.exitCode = 10;
|
|
116
|
+
}
|
|
117
|
+
} catch (error) {
|
|
118
|
+
console.error(`PR digest incomplete: ${error.message}`);
|
|
119
|
+
process.exitCode = 2;
|
|
120
|
+
}
|
|
@@ -0,0 +1,102 @@
|
|
|
1
|
+
import { readFileSync } from 'node:fs';
|
|
2
|
+
|
|
3
|
+
const bindings = { codex: 'codex', claude: 'claudeAgent', grok: 'grok', antigravity: 'antigravity' };
|
|
4
|
+
const effortIds = { codex: 'reasoningEffort', claude: 'effort', grok: 'reasoningEffort' };
|
|
5
|
+
const nonempty = (value) => typeof value === 'string' && value.trim().length > 0;
|
|
6
|
+
|
|
7
|
+
function argumentsFrom(args) {
|
|
8
|
+
const options = { excluded: [] };
|
|
9
|
+
for (let i = 0; i < args.length; i += 2) {
|
|
10
|
+
const flag = args[i];
|
|
11
|
+
if (!['--provider', '--capabilities', '--class', '--model', '--effort', '--exclude'].includes(flag)) {
|
|
12
|
+
throw new Error(`unexpected argument ${flag}`);
|
|
13
|
+
}
|
|
14
|
+
const value = args[i + 1];
|
|
15
|
+
if (!nonempty(value) || value.startsWith('--')) throw new Error(`expected value after ${flag}`);
|
|
16
|
+
if (flag === '--exclude') options.excluded.push(value);
|
|
17
|
+
else {
|
|
18
|
+
const key = flag.slice(2);
|
|
19
|
+
if (key in options) throw new Error(`duplicate argument ${flag}`);
|
|
20
|
+
options[key] = value;
|
|
21
|
+
}
|
|
22
|
+
}
|
|
23
|
+
if (!options.provider || !options.capabilities || !options.effort) {
|
|
24
|
+
throw new Error('expected --provider codex|claude|grok|antigravity --capabilities path --effort level');
|
|
25
|
+
}
|
|
26
|
+
return options;
|
|
27
|
+
}
|
|
28
|
+
|
|
29
|
+
function providerFrom(catalog, provider) {
|
|
30
|
+
if (!Array.isArray(catalog?.providers) || catalog.providers.some((p) =>
|
|
31
|
+
!p || !nonempty(p.providerInstanceId) || !nonempty(p.driverKind))) {
|
|
32
|
+
throw new Error('malformed capabilities providers');
|
|
33
|
+
}
|
|
34
|
+
const matches = catalog.providers.filter((p) => p.providerInstanceId === bindings[provider]);
|
|
35
|
+
if (matches.length !== 1 || matches[0].driverKind !== bindings[provider]) {
|
|
36
|
+
throw new Error(`unknown or ambiguous provider instance ${bindings[provider]}`);
|
|
37
|
+
}
|
|
38
|
+
const instance = matches[0];
|
|
39
|
+
if (!Array.isArray(instance.models) || instance.models.some((model) =>
|
|
40
|
+
!model || !nonempty(model.id) || !Array.isArray(model.options) || model.options.some((option) =>
|
|
41
|
+
!option || !nonempty(option.id) || (option.options !== undefined &&
|
|
42
|
+
(!Array.isArray(option.options) || option.options.some((value) => !value || !nonempty(value.id))))))) {
|
|
43
|
+
throw new Error('malformed capabilities models or options');
|
|
44
|
+
}
|
|
45
|
+
return instance;
|
|
46
|
+
}
|
|
47
|
+
|
|
48
|
+
function modelFrom(models, { provider, model: pin, class: modelClass, excluded }) {
|
|
49
|
+
if (pin && pin !== 'null') {
|
|
50
|
+
const model = models.find((model) => model.id === pin && !excluded.includes(model.id));
|
|
51
|
+
if (!model) throw new Error(`missing requested model ${pin}`);
|
|
52
|
+
return model;
|
|
53
|
+
}
|
|
54
|
+
if (!pin && !modelClass) throw new Error('expected --class or --model (use --model null for a null role)');
|
|
55
|
+
if (!modelClass) {
|
|
56
|
+
if (['codex', 'claude'].includes(provider)) {
|
|
57
|
+
throw new Error(`intentional absence for ${provider}: missing model and class`);
|
|
58
|
+
}
|
|
59
|
+
// Launch-by-agent-ID providers bind the first listed ID; exclusions never select a substitute.
|
|
60
|
+
const model = models[0];
|
|
61
|
+
if (!model || excluded.includes(model.id)) throw new Error('missing first listed model');
|
|
62
|
+
return model;
|
|
63
|
+
}
|
|
64
|
+
if (!['codex', 'claude'].includes(provider) || !/^[a-z]+$/.test(modelClass)) {
|
|
65
|
+
throw new Error(`unsupported model class ${modelClass} for ${provider}`);
|
|
66
|
+
}
|
|
67
|
+
const pattern = provider === 'codex'
|
|
68
|
+
? new RegExp(`^gpt-(\\d+(?:\\.\\d+)*)-${modelClass}$`)
|
|
69
|
+
: new RegExp(`^claude-${modelClass}-(\\d+-\\d+)$`);
|
|
70
|
+
const candidates = models.filter((model) => pattern.test(model.id) && !excluded.includes(model.id));
|
|
71
|
+
candidates.sort((a, b) => {
|
|
72
|
+
const left = a.id.match(pattern)[1].split(/[.-]/).map(Number);
|
|
73
|
+
const right = b.id.match(pattern)[1].split(/[.-]/).map(Number);
|
|
74
|
+
for (let i = 0; i < Math.max(left.length, right.length); i++) {
|
|
75
|
+
const difference = (right[i] ?? 0) - (left[i] ?? 0);
|
|
76
|
+
if (difference) return difference;
|
|
77
|
+
}
|
|
78
|
+
return 0;
|
|
79
|
+
});
|
|
80
|
+
if (!candidates.length) throw new Error(`no matching ${modelClass} model for ${provider}`);
|
|
81
|
+
return candidates[0];
|
|
82
|
+
}
|
|
83
|
+
|
|
84
|
+
try {
|
|
85
|
+
const options = argumentsFrom(process.argv.slice(2));
|
|
86
|
+
const { provider, capabilities: path, effort } = options;
|
|
87
|
+
if (!Object.hasOwn(bindings, provider)) throw new Error(`unknown provider ${provider}`);
|
|
88
|
+
const catalog = JSON.parse(readFileSync(path, 'utf8'));
|
|
89
|
+
const instance = providerFrom(catalog, provider);
|
|
90
|
+
const model = modelFrom(instance.models, options);
|
|
91
|
+
const effortOptions = model.options.filter((option) => effortIds[provider]
|
|
92
|
+
? option.id === effortIds[provider] : ['reasoningEffort', 'effort'].includes(option.id));
|
|
93
|
+
if (effortOptions.length !== 1 || !effortOptions[0].options?.some((value) => value.id === effort)
|
|
94
|
+
|| (provider === 'grok' && effort === 'max')) {
|
|
95
|
+
throw new Error(`unsupported effort ${effort} for ${provider}/${model.id}`);
|
|
96
|
+
}
|
|
97
|
+
console.log(JSON.stringify({ provider, providerInstanceId: instance.providerInstanceId,
|
|
98
|
+
model: model.id, effortOption: { id: effortOptions[0].id, value: effort }, source: 'capabilities', path }));
|
|
99
|
+
} catch (error) {
|
|
100
|
+
console.error(`model capabilities resolution hold: ${error.message}`);
|
|
101
|
+
process.exitCode = 1;
|
|
102
|
+
}
|
|
@@ -5,6 +5,9 @@ description: When exploring or planning engineering work, use axstack-align to s
|
|
|
5
5
|
|
|
6
6
|
# Align
|
|
7
7
|
|
|
8
|
+
For authorized delivery runs, follow [Autopilot](../axstack/references/autopilot.md)
|
|
9
|
+
for phase continuation and holds.
|
|
10
|
+
|
|
8
11
|
On driver entry, sweep under [Workspace hygiene](../axstack/references/workspace-hygiene.md); dispatched workers do not sweep.
|
|
9
12
|
For every dispatch brief, name its private `<run dir>/evidence/<dispatch>/` folder.
|
|
10
13
|
|
|
@@ -34,7 +37,7 @@ substantial; apply routing's existing size reassessment rule.
|
|
|
34
37
|
tools before asking the user. Separate facts from preferences, name evidence
|
|
35
38
|
gaps, and map which decisions unlock others. When a fact needed for the
|
|
36
39
|
frontier is not derivable from the local repo or docs by ordinary reading,
|
|
37
|
-
dispatch `axstack-research` branches through
|
|
40
|
+
dispatch `axstack-research` branches through T3 `delegate_task` by source type:
|
|
38
41
|
requirements, code, web, and, once configured, X. Give one owner per branch,
|
|
39
42
|
use cross-harness routes where the roles allow, and require a cited note per
|
|
40
43
|
the research skill's source standards. The driver folds verified claims into
|
|
@@ -73,12 +76,17 @@ material disagreement remains, then surface the choices to the user. Never
|
|
|
73
76
|
fabricate consensus or impersonate a role.
|
|
74
77
|
|
|
75
78
|
Immediately before the first actual adviser dispatch, load and follow
|
|
76
|
-
[
|
|
79
|
+
[T3 runtime](../axstack/references/t3-runtime.md). Reuse each adviser
|
|
77
80
|
session and settled receipt; consult only the changed frontier and reuse
|
|
78
81
|
unchanged receipts. Record compact adviser evidence, the driver's assessment,
|
|
79
82
|
and user-resolved choices for `axstack-spec`. If either adviser is unavailable,
|
|
80
83
|
hold Align; safe fact work may continue without substitution.
|
|
81
84
|
|
|
85
|
+
An optional adviser note may be deferred or rejected in a `Decisions` row with
|
|
86
|
+
the draft unchanged; it needs no new adviser pair. Changed draft text, a
|
|
87
|
+
blocking finding, or a high-stakes decision requires fresh receipts on the new
|
|
88
|
+
revision.
|
|
89
|
+
|
|
82
90
|
## Arena for hard-to-reverse design choices
|
|
83
91
|
|
|
84
92
|
Critique of one draft anchors every reader to that draft's shape. Rung 2 designs
|
|
@@ -122,10 +130,15 @@ an arena. Small or routine questions never enter the arena.
|
|
|
122
130
|
Record the synthesis note (base, grafts and their source candidate, rejections,
|
|
123
131
|
dropouts, judge verdicts per round) as `Decisions` rows in the
|
|
124
132
|
[run record](../axstack/references/run-record.md). Load
|
|
125
|
-
[
|
|
126
|
-
first candidate or judge dispatch. If
|
|
127
|
-
|
|
128
|
-
|
|
133
|
+
[T3 runtime](../axstack/references/t3-runtime.md) immediately before the
|
|
134
|
+
first candidate or judge dispatch. If an optional Grok or Antigravity candidate
|
|
135
|
+
malfunctions (launch failure, trust/login prompt, or prompt block), fence it,
|
|
136
|
+
record `absent (<reason>)`, name it once in the next
|
|
137
|
+
read-back, and continue with available candidates without relay or substitution.
|
|
138
|
+
A required adviser, candidate, or judge unavailable at launch or returning a
|
|
139
|
+
failed receipt holds that question without substitution; record the gap and ask
|
|
140
|
+
whether to proceed. In mixed fan-out retain at least one Codex and one Claude
|
|
141
|
+
seat, or hold the affected question.
|
|
129
142
|
For an uncertain dispatch, reconcile natively; it is never treated as absent.
|
|
130
143
|
Unaffected fact work and questions continue.
|
|
131
144
|
|
|
@@ -166,7 +179,7 @@ record or spec. Read-only scope keeps proposed documentation in the permitted
|
|
|
166
179
|
private record or response. Documentation is neither implementation nor spec
|
|
167
180
|
approval; record chosen document names and paths once per run.
|
|
168
181
|
|
|
169
|
-
## Read back, classify, and
|
|
182
|
+
## Read back, classify, and route
|
|
170
183
|
|
|
171
184
|
1. Read back the decisions, constraints, exclusions, and remaining evidence
|
|
172
185
|
gaps. For substantial work, this summary becomes part of the draft spec in
|
|
@@ -186,9 +199,10 @@ approval; record chosen document names and paths once per run.
|
|
|
186
199
|
documentation pointers without adding another runtime. Return the compact
|
|
187
200
|
scope and record pointer in the current chat. Native transfer is separate:
|
|
188
201
|
use it only when the user explicitly requests transfer, loading
|
|
189
|
-
[
|
|
202
|
+
[T3 runtime](../axstack/references/t3-runtime.md) immediately before
|
|
190
203
|
actual dispatch. Alignment completion never dispatches a recipient.
|
|
191
204
|
|
|
192
|
-
Alignment
|
|
193
|
-
identity is explicit
|
|
194
|
-
|
|
205
|
+
Alignment completes for both sizes only when the handoff is usable and its next
|
|
206
|
+
scope identity is explicit. An eligible delivery run continues under Autopilot;
|
|
207
|
+
an explicit stop-after-Align request ends here. Substantial work continues to
|
|
208
|
+
Spec, and small work continues from its small-change intent to Implement.
|