axstack 0.14.2 → 0.16.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -12,9 +12,9 @@ role choices, evidence, review policy, and one private derived run record.
12
12
 
13
13
  ## How a run works
14
14
 
15
- Invoke `axstack` or the needed phase directly: `axstack-align`,
16
- `axstack-spec`, `axstack-tickets`, `axstack-implement`, `axstack-review`, and
17
- `axstack-watch`. Direct `axstack-research`, `axstack-explain`,
15
+ Invoke the needed phase directly: `axstack-align`, `axstack-spec`,
16
+ `axstack-tickets`, `axstack-implement`, `axstack-review`, and `axstack-watch`.
17
+ Direct `axstack-research`, `axstack-explain`,
18
18
  `axstack-improve`, and `axstack-debug` routes need no spec ceremony. `axstack-relay` remains an
19
19
  optional inline route for explicit messages and authorized notifications; an
20
20
  unavailable or legacy-runtime-only relay falls back to the current conversation
@@ -68,8 +68,8 @@ See [installation details](docs/installation.md).
68
68
 
69
69
  `--instructions` manages one versioned Axstack block in `AGENTS.md`,
70
70
  `CLAUDE.md`, or `GEMINI.md`. Harness defaults resolve those files automatically. The block
71
- points at the installed entry skill and requires every subagent, delegated worker,
72
- reviewer, and cross-harness dispatch to use visible Orca orchestration via the `orca` CLI
71
+ requires direct matching phase-skill invocation and requires every subagent, delegated
72
+ worker, reviewer, and cross-harness dispatch to use visible Orca orchestration via the `orca` CLI
73
73
  rather than a harness-native subagent tool (e.g. Claude/Codex native subagents). OpenCode
74
74
  and Antigravity subagents run as Orca-supervised workers. Text and file
75
75
  mode outside the markers are preserved; edited, malformed, unowned, or unsafe
@@ -147,7 +147,7 @@ to rewrite them.
147
147
 
148
148
  ## Role behavior after installation
149
149
 
150
- The runtime reads `roles.json` relative to the actually loaded `axstack` skill.
150
+ The runtime reads `roles.json` from the installed shared root `skills/axstack/`.
151
151
  A new run records the selected preset plus all 23 role rows. An active run keeps
152
152
  that snapshot after a later preset install unless the user explicitly changes
153
153
  it and accepts the resulting evidence invalidation.
package/docs/workflows.md CHANGED
@@ -10,9 +10,9 @@ Axstack implements that skill's specialist capability.
10
10
 
11
11
  ## Routing and scope identity
12
12
 
13
- `axstack` classifies the request and loads only the applicable phase plus shared
14
- references for routing, lifecycle, Orca runtime boundaries, role/model/risk
15
- contracts, the run record, and PR shape.
13
+ The directly invoked phase loads the applicable shared references for routing,
14
+ lifecycle, Orca runtime boundaries, role/model/risk contracts, the run record,
15
+ and PR shape.
16
16
 
17
17
  Direct routes need no spec ceremony:
18
18
 
@@ -53,7 +53,7 @@ role.
53
53
 
54
54
  The installed `<skills-dir>/axstack/roles.json` adds the selected preset name:
55
55
  `{ "version": 1, "preset": "<name>", "roles": [...] }`. The runtime reads it
56
- relative to the actually loaded `axstack` skill and records the whole table for
56
+ from the installed shared root `skills/axstack/` and records the whole table for
57
57
  a new run. Active runs retain their snapshot after later installation changes.
58
58
 
59
59
  Peer roles keep the stable IDs `axstack-reviewer-primary` and
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "axstack",
3
- "version": "0.14.2",
3
+ "version": "0.16.0",
4
4
  "description": "Axstack installer and setup CLI: installs owned chat skills and role data, configures supported harness settings, and checks Orca capabilities.",
5
5
  "keywords": [
6
6
  "claude-code",
@@ -2,7 +2,7 @@
2
2
 
3
3
  Read this when the current session is the native Orca PR **driver** or its
4
4
  **watchdog**. The approved contract is
5
- `docs/specs/pr-automations.md` revision 5; this reference restates the parts an
5
+ `docs/specs/pr-automations.md` revision 6; this reference restates the parts an
6
6
  automation session must execute and does not widen them.
7
7
 
8
8
  Pair A/B is retired for this contract; its artefacts remain untouched.
@@ -11,7 +11,7 @@ Pair A/B is retired for this contract; its artefacts remain untouched.
11
11
 
12
12
  - **driver** — every 15 minutes, fresh Opus session in the host's `root`
13
13
  folder workspace, which is not a git repository and belongs to no project. It discovers GitHub work, binds the persistent Orca
14
- Run, checks each selected PR's head out into a free project-local slot, dispatches the
14
+ Run, creates one worktree per selected PR and checks its head out there, dispatches the
15
15
  matching Axstack agent, reconciles completions and decisions, then exits. The
16
16
  driver is the automation session itself, with no `axstack-monitor` or
17
17
  `axstack-owner` role row. It never performs review or repair in its own
@@ -97,7 +97,7 @@ The driver performs this order and exits:
97
97
  inbox. For a `worker_done` matching a live marker, verify the review id at
98
98
  the bound head, push range, or opened token. Release the worker; once its
99
99
  release receipt is settled and the process has exited, run the cleanup
100
- under "Run directory" below, which proves the slot disposable before
100
+ under "Run directory" below, which proves the worktree disposable before
101
101
  resetting it; then clear the marker. Unverifiable delivery stays in
102
102
  `pending_settlement[]` and blocks only that PR.
103
103
  3. Consume decisions as their sole consumer under "Decision tokens" below.
@@ -126,34 +126,38 @@ The driver performs this order and exits:
126
126
 
127
127
  Claude Code trusts a folder per git toplevel and stops at its "Quick safety
128
128
  check" dialog otherwise, and the driver never answers that dialog for a
129
- worker. So workers run in a fixed pool: every allowlisted project has
130
- exactly five slot worktrees, `slot-1` through `slot-5`, its Orca child
131
- worktrees created once, parented to the project's primary worktree, and trusted once
132
- by the user through that dialog. The driver never creates or removes a
133
- worktree and never writes `~/.claude.json`. A slot is free when no live
134
- dispatch marker names it, it is not in `retained_slots[]`, no terminal is
135
- listed in it, and its tree is clean; no free slot in the project defers the
136
- PR to `deferred[]` like a budget, not a hold. Taking a slot fetches the head into the project clone,
137
- checks the slot out detached at the pinned head, and verifies HEAD equals
138
- it; the marker's worktree is the slot path. The worker is launched by Orca
139
- itself — `worker-start --agent claude --model claude-opus-5 --effort medium`
140
- on that slot — so the runtime owns the process: `worker-release` ends it and
141
- `worker-show` proves it exited, which is what returns the slot to the pool.
142
- Never pre-create the worker's terminal or hand a terminal handle to
129
+ worker. Trust inherits from the project's primary clone, which the user
130
+ trusted once (a fresh child worktree launched with no dialog — canary
131
+ 2026-09-18). There is no fixed pool: fetch the head into the project clone first (a
132
+ failed fetch is a health line and no dispatch; no worktree exists yet), then
133
+ create one Orca worktree per dispatch, `orca worktree create --repo id:<clone
134
+ id> --name <repo short>-<num>-<head7> --base-branch <default branch>
135
+ --parent-worktree id:<clone id>::<clone path> --setup skip` (a name collision
136
+ with a retained worktree at the same head appends the tick's
137
+ `tick_started_at` stamp); close the creation terminal Orca opens in it with
138
+ `orca terminal close --worktree <selector> --all`; check the worktree out
139
+ detached at the pinned head and verify HEAD equals it; the marker's worktree
140
+ is that path. Never write
141
+ `~/.claude.json`. The worker is launched by Orca itself — `worker-start
142
+ --agent claude --model claude-opus-5 --effort medium` in that worktree — so
143
+ the runtime owns the process: `worker-release` ends it and `worker-show`
144
+ proves it exited, which is what allows the worktree to be removed. Never
145
+ pre-create the worker's terminal or hand a terminal handle to
143
146
  `worker-start`: a reused handle is a resource Orca labels `external`, one it
144
- can neither stop nor prove exited, so every such slot ends retained. After
145
- `worker-start` run `worker-show` on the receipt's dispatch id and require
146
- `projection.resource.state == owned`; anything else is a launch Orca does not
147
- own: apply the runtime-refusal recovery rules under "Safety holds" (the
148
- `worker-list` row's `nextAction` argv verbatim; `none` means inspect and
149
- retain), append a `worker not owned` health line and a `retained_slots[]`
150
- entry for the slot, defer the PR, and record no marker. Orca's per-agent default arguments supply
151
- `--dangerously-skip-permissions`; the brief loads the skill files it needs by
152
- path. If `worker-start` reports a failed stage or a visible hold (the "Quick
153
- safety check" trust dialog) the slot is not trusted: name the slot in a health
154
- line, defer the PR, dispatch nothing, never answer the dialog. Project
155
- customizations load as they would for the user; the allowlist is defi-com
156
- only and the user accepted that surface on 2026-09-18.
147
+ can neither stop nor prove exited, so every such worktree ends retained.
148
+ After `worker-start` run `worker-show` on the receipt's dispatch id and
149
+ require `projection.resource.state == owned`; anything else is a launch Orca
150
+ does not own: apply the runtime-refusal recovery rules under "Safety holds"
151
+ (the `worker-list` row's `nextAction` argv verbatim; `none` means inspect
152
+ and retain), append a `worker not owned` health line and a `retained_slots[]`
153
+ entry for the worktree, defer the PR, and record no marker. Orca's per-agent
154
+ default arguments supply `--dangerously-skip-permissions`; the brief loads
155
+ the skill files it needs by path. If `worker-start` reports a failed stage
156
+ or a visible hold (the "Quick safety check" trust dialog) the worktree is not
157
+ trusted: name it in a health line, defer the PR, dispatch nothing, never
158
+ answer the dialog, and remove the worktree through the no-worker branch.
159
+ Project customizations load as they would for the user; the allowlist is
160
+ defi-com only and the user accepted that surface on 2026-09-18.
157
161
 
158
162
  Every selected PR receives one dispatch marker with task id, dispatch id,
159
163
  worktree, head, `started_at`, reservation (`verdict` or `repair`), and trigger:
@@ -175,16 +179,22 @@ An own PR needs repair when either trigger applies:
175
179
  digest not recorded for that PR and head. Both keys are required: the same
176
180
  finding under a new review id must not re-trigger repair. Record review id
177
181
  and body digest when dispatching. A superseded head with a new review
178
- triggers again subject to the 24 h cap.
179
-
180
- Repair also requires no deploy-on-push head branch, no live repair cap, and
181
- selection of the lowest own PR in its stack that needs repair. Take a free
182
- slot of that project (its slots are parented to that project's primary
183
- worktree, so the work appears under the project it serves) at the exact
184
- head, and dispatch one `axstack-watch` agent in authored repair mode. Its
182
+ triggers again at the new head.
183
+
184
+ Repair also requires no deploy-on-push head branch, a head not already in
185
+ `repaired_heads[]`, and selection of the lowest own PR in its stack that
186
+ needs repair. Create the
187
+ dispatch worktree (parented to that project's primary worktree, so the work
188
+ appears under the project it serves) at the exact head, and dispatch one
189
+ `axstack-watch` agent in authored repair mode. Its
185
190
  brief contains only the triggering checks or review findings. Each open
186
191
  descendant records one user-owned `pending restack` hold until it stops needing
187
- repair. The 24 h cap starts at dispatch and an abandon does not refund it.
192
+ repair. At dispatch record the head in `repaired_heads[]` (`pr`, `head`,
193
+ `dispatched_at`): one repair per head, no time cap. A repair pushes a new
194
+ head; a still-not-merge-ready new head shows a new failing check or review
195
+ and is repaired again; a repair that pushes nothing is not retried at that
196
+ head until a human or a new commit moves it. A confirmed abandon removes the
197
+ record (retry once).
188
198
 
189
199
  A debounced peer PR is eligible when self has not reviewed its head. Read
190
200
  `gh pr view --json reviews` before dispatch. Whenever any self review with
@@ -196,14 +206,15 @@ human-placed block is never overwritten. A dismissed block and a self-approved
196
206
  PR are skipped. Dispatch one `axstack-review` agent in peer mode and link the
197
207
  prior review in its brief.
198
208
 
199
- Budgets are one `verdict` dispatch per tick, oldest first; at most six `repair`
200
- markers live across all repositories; and one repair per PR per 24 h. Put every
201
- eligible PR not dispatched because of a budget in `deferred[]` with repo, PR,
202
- and head. Budget exhaustion records the count and is not a hold.
209
+ The only concurrency limit is the host-wide cap: at most eight live dispatch
210
+ markers across all repositories and both reservations, oldest eligible
211
+ first; plus one repair per head. Put every eligible PR not dispatched
212
+ because the cap is reached in `deferred[]` with repo, PR, and head. Reaching
213
+ the cap records the count and is not a hold.
203
214
 
204
215
  ## Agents and verdicts
205
216
 
206
- Every agent works in a project-local slot worktree checked out detached at
217
+ Every agent works in its own project-local worktree checked out detached at
207
218
  the exact head and reports only through the Orca worker protocol.
208
219
 
209
220
  Peer review runs the two isolated configured reviewers on the identical brief,
@@ -343,12 +354,12 @@ is no gate for health findings.
343
354
  ## Run directory
344
355
 
345
356
  The driver and watchdog run from the host's `root` folder workspace, not a
346
- project worktree: no project owns the automation, and every slot worktree
347
- belongs to the project it serves. That workspace is not a git repository, so
357
+ project worktree: no project owns the automation, and every dispatch
358
+ worktree belongs to the project it serves. That workspace is not a git repository, so
348
359
  the run directory is private host state, one
349
360
  `~/.local/share/axstack/runs/<run id>/` directory. Settlement leaves nothing
350
361
  behind, but never destroys work. Before any destructive step the driver
351
- proves the slot is disposable: the worker is settled — on the release path
362
+ proves the worktree is disposable: the worker is settled — on the release path
352
363
  a settled release receipt, on the abandon path an accepted abandon receipt,
353
364
  either with proven process exit; pending or unknown stops here — the
354
365
  worktree's HEAD is either the pinned head or a candidate that is durably
@@ -358,20 +369,21 @@ targeted fetch of that exact remote branch into a per-dispatch ref, never
358
369
  shared clone can fake durability, a failed fetch retaining the worktree — or held by a
359
370
  `refs/axstack/decisions/<token>` ref in the project clone; and, on the
360
371
  abandon path, the worktree has no uncommitted changes.
361
- Only then it closes any terminal tab still listed, resets the slot to the
362
- pinned head, clears untracked artefacts, and verifies the slot is clean, so
363
- it is back in the pool. On the settled path
364
- the worker has finished, so untracked files are artefacts by definition and
365
- are cleared; a candidate there is already pushed or token-held. A slot
366
- whose release is settled but whose HEAD cannot be proven disposable, and an
367
- abandoned slot that is dirty or holds an unproven candidate, are both
368
- **retained**: named in one `health[]` line with path and SHA, and blocking
369
- only that PR with the user as owner. Retention is mechanical on both paths:
370
- the driver appends `retained_slots[]` `{slot, pr, head, reason}`, which the
371
- free predicate excludes for every PR — a clean, terminal-less slot holding
372
- an unpushed candidate would otherwise look free — and an entry is cleared
373
- only by the user after reconciling the candidate. A dirty slot or a terminal that
374
- outlives its dispatch without such a retention record is a health finding. The run directory contains:
372
+ Only then it closes any terminal tab still listed and removes the worktree
373
+ with `orca worktree rm --worktree <selector> --force` (the proof is the
374
+ gate; a finished worker's untracked artefacts are not), so the dispatch
375
+ leaves nothing behind. On the settled path the worker has finished, so untracked
376
+ files are artefacts by definition; a candidate there is already pushed or
377
+ token-held. A worktree whose release is settled but whose HEAD cannot be
378
+ proven disposable, and an abandoned worktree that is dirty or holds an
379
+ unproven candidate, are both **retained** — retained in place: named in one `health[]`
380
+ line with path and SHA, and blocking only that PR with the user as owner.
381
+ Retention is mechanical on both paths: the driver appends `retained_slots[]`
382
+ `{slot, pr, head, reason}` (`slot` is the worktree path); a retained worktree
383
+ is never removed by the driver, and the entry is cleared only by the user
384
+ after reconciling the candidate, who also removes the worktree. A leftover
385
+ worktree or terminal that outlives its dispatch without such a retention
386
+ record is a health finding. The run directory contains:
375
387
 
376
388
  - `cursor.json` — driver only, with these exact keys: `fingerprint`,
377
389
  `tick_started_at`, `tick_done_at`, `tick_outcome`, `prs{url: {head, base,
@@ -379,7 +391,10 @@ outlives its dispatch without such a retention record is a health finding. The r
379
391
  `task_id`, `dispatch_id`, `worktree`, `head`, `started_at`, `reservation`,
380
392
  `trigger`), `deferred[]`, `pending_settlement[]`,
381
393
  `retained_slots[]` (`slot`, `pr`, `head`, `reason`),
382
- `repair_caps{url: {expires_at}}`, `abandon_count{head: n}`,
394
+ `repaired_heads[]` (`pr`, `head`, `dispatched_at`),
395
+ `repair_caps{url: {expires_at}}` (legacy: the first rev-6 tick clears it to
396
+ `{}` with one health line; never written again),
397
+ `abandon_count{head: n}`,
383
398
  `processed_reviews[]` (`review_id`, `pr`, `head`, `digest`),
384
399
  `deploy_on_push{repo: [branches]}`,
385
400
  `legacy_automation_reviews[]`, `health[]`,
@@ -420,12 +435,12 @@ delivery uses [axstack-relay](../../axstack-relay/SKILL.md).
420
435
  `runtime_refusal {code, first_seen, last_seen}`; the same code keeps the
421
436
  hold and updates `last_seen` without a new health line, a different code is
422
437
  a new finding. Prose is never the key. When `worker-start` itself is
423
- refused after the slot was checked out, there is no worker, so the
438
+ refused after the worktree was created, there is no worker, so the
424
439
  settlement proof does not apply; the driver reads the receipt's `failedStage`
425
440
  and `residualResources` first. With no Dispatch and no residual resources
426
441
  the no-worker branch applies: the worktree's HEAD must equal the pinned
427
- head and `git status --porcelain` must be empty, and then the slot is simply
428
- free again in the same tick; there is nothing to remove.
442
+ head and `git status --porcelain` must be empty, and then the worktree is
443
+ simply removed in the same tick.
429
444
  With a Dispatch or any residual resource the failed start owns runtime
430
445
  state, and retaining alone is not recovery: the driver follows the
431
446
  runtime's recovery guide. With a Dispatch: `worker-list` for that run, and
@@ -4,11 +4,11 @@ Apply these authority, scope, and model rules before consequential action.
4
4
 
5
5
  ## Required lifecycle load
6
6
 
7
- Except for `axstack-audit` itself, every independently called phase must load
8
- and follow [Shared lifecycle](lifecycle.md) before acting. When a substantive
9
- run ends or reaches a meaningful checkpoint, apply the lifecycle audit hook.
10
- The audit phase loads these contracts, writes its assigned record, and stops;
11
- it never audits itself.
7
+ Except for `axstack-audit` and `axstack-relay`, every independently called phase
8
+ must load and follow [Shared lifecycle](lifecycle.md) before acting. When a
9
+ substantive run ends or reaches a meaningful checkpoint, apply the lifecycle
10
+ audit hook. The audit phase loads these contracts, writes its assigned record,
11
+ and stops; it never audits itself.
12
12
 
13
13
  ## Scope identity (conditional — see routing and lifecycle)
14
14
 
@@ -1,25 +1,23 @@
1
1
  # Shared lifecycle and receipts
2
2
 
3
- Phases load through [Standing contracts](contracts.md)' mandatory edge.
4
- Substantive delegated or resumable work uses a driver-owned [Run record](run-record.md)
3
+ Phases load through [Standing contracts](contracts.md).
4
+ Delegated or resumable work uses a driver-owned [Run record](run-record.md)
5
5
  binding state and receipts to exact revisions.
6
6
 
7
7
  ## Roster (compact)
8
8
 
9
9
  - Driver: current chat; owns scope, decisions, cross-PR dependencies, Linear
10
10
  mutations, and integration.
11
- - Owner (`axstack-owner`): one persistent owner per PR; launches its author,
12
- reviewers and watches; may perform authorized PR-scoped publication within user authority.
13
- Human merge is default.
11
+ - Owner: driver owns loop PRs; `axstack-owner` only for standalone watch/review
12
+ without live driver. It may perform PR-scoped publication within user authority. Human
13
+ merge is default.
14
14
  - Author: exactly one writer per candidate; accepted fixes return there.
15
15
  Workers launch no recursive teams.
16
16
  - Reviewers: peer = two independent `axstack-reviewer-primary` and
17
17
  `axstack-reviewer-secondary` sessions with identical brief and isolated first
18
18
  pass; authored = one eligible configured reviewer from actual author
19
19
  provenance. Owner and author never review.
20
- - Driver/monitor/watchdog: the driver every 15 minutes dispatches and exits as
21
- a mutating owner; the watchdog is model-free and read-only, has no gate, and
22
- records `watchdog.log`; there is no watch deadline for automations.
20
+ - Automation driver/monitor/watchdog: see [Watch health](#watch-health).
23
21
  - Auditor (`axstack-auditor`): report-only; never edits, merges, activates, or
24
22
  audits itself.
25
23
 
@@ -78,10 +76,17 @@ Store concise receipt references, not raw worker output, in the [Run record](run
78
76
  ## Execution tracking
79
77
 
80
78
  The driver consumes native Orca completion and escalation deliveries for the
81
- active Run. Process each whole delivery before acknowledgment and validate its
82
- Task, Dispatch, sender, authority, revisions, and receipts before advancing the
83
- run record. Duplicate deliveries are deduplicated by runtime identity. Healthy
84
- unchanged observations produce no user-facing update.
79
+ active Run. A driver turn does not end while a Dispatch is unsettled unless one
80
+ completion wait from the orchestration guide is armed (background where the
81
+ harness supports it, foreground otherwise) and re-armed on timeout; sleep or
82
+ poll loops are forbidden. An explicitly invoked phase dispatches its configured
83
+ roles through Orca and closes with the lifecycle close-out; in-chat execution
84
+ covers only ordinary reading, writing, and local checks. Heartbeat deliveries
85
+ are acknowledged with no user-facing text. Process each whole delivery before
86
+ acknowledgment and validate its Task, Dispatch, sender, authority, revisions,
87
+ and receipts before advancing the run record. Duplicate deliveries are
88
+ deduplicated by runtime identity. Healthy unchanged observations produce no
89
+ user-facing update.
85
90
 
86
91
  Detect completed-but-unadvanced work, failed sessions, unresolved launch
87
92
  receipts, and stalls through the version-matched orchestration guide. Never
@@ -98,7 +103,9 @@ Tracking grants no merge, release, model-substitution, or scope authority.
98
103
 
99
104
  The default 24-hour deadline covers standalone task-owned timers. Stop them at
100
105
  deadline and preserve remaining work; there is no watch deadline for
101
- automations. Merge-ready differs from merged; human merges.
106
+ automations. A PR is merge-ready only with the applicable review receipt(s) at
107
+ its exact head; green CI or tests alone never make it merge-ready. Merge-ready
108
+ differs from merged; human merges.
102
109
 
103
110
  ## Watch health
104
111
 
@@ -110,24 +117,24 @@ no watch deadline for automations. Build no custom scheduler and use no legacy
110
117
  fallback. Details live in
111
118
  [Watch runtime](../../axstack-watch/references/watch-runtime.md).
112
119
 
113
- ## Audit hook (end of run and meaningful checkpoints)
114
-
115
- Auditing defaults on for every substantive run at its end and meaningful
116
- checkpoints such as material deviation or repeated repair. Load the bundled [audit skill](../../axstack-audit/SKILL.md)
117
- and dispatch its auditor. An `axstack-audit` run is excluded: it writes its
118
- record and launches no children.
119
-
120
- The auditor reads the [Run record](run-record.md) for scope, outcomes, and
121
- metric counts/denominators; reports evidenced PASS/FAIL/UNKNOWN; and invents no
122
- numbers or cost. Proposals change nothing. Accepted proposals return as
123
- tested, independently reviewed work with a regression scenario and unchanged
124
- holdout checks. No automatic self-edit, merge, or activation. Records stay
125
- private; publication needs separate authority.
126
-
127
- ## Idle-complete archive and retain
128
-
129
- After required PRs merge or hand off, timers stop, and receipts verify, mark
130
- the same [Run record](run-record.md) `Archived` in place. Preserve
131
- scope, revisions, evidence, receipts, and expiries. Archive only idle-complete
132
- records: never active/waiting workers, unrelated host state, or ownership merely
133
- because it is idle.
120
+ ## Audit hook (close-out and meaningful checkpoints)
121
+
122
+ Audit measurement is enabled by default for every substantive run; dispatch
123
+ follows [Close-out](#close-out) or a material-deviation/repeated-repair
124
+ checkpoint. An `axstack-audit` run is excluded; it launches no children.
125
+ Load [axstack-audit](../../axstack-audit/SKILL.md). Accepted proposals
126
+ require a regression scenario and unchanged holdout checks; they change
127
+ nothing without tested independent review.
128
+
129
+ ## Close-out
130
+
131
+ After required PRs merge by forge state—not local branch ancestry—close out in
132
+ order: (1) settle every worker terminal through the orchestration guide; (2)
133
+ write a compact record with counts and denominators for
134
+ user interventions, deviations from plan, and repairs; (3) dispatch
135
+ `axstack-auditor` only when any count is non-zero or the user asks—an unavailable
136
+ auditor leaves close-out pending, never skipped silently; (4)
137
+ release merged run worktrees and branches and close Linear tickets
138
+ (driver-owned); (5) mark the [Run record](run-record.md) `Archived`. The driver
139
+ cannot report the run done before steps (1)-(5) have receipts. Small one-step
140
+ lookups keep the run record's exemption.
@@ -22,7 +22,7 @@ daemon, scheduler, database, or escalation engine.
22
22
 
23
23
  ## Bind the configured role
24
24
 
25
- Read `roles.json` relative to the actually loaded `axstack` skill. The installed
25
+ Read `roles.json` from the installed shared root `skills/axstack/`. The installed
26
26
  shape is `{ "version": 1, "preset": "<name>", "roles": [...] }`. Bundled
27
27
  profiles are setup inputs shaped as
28
28
  `{ "version": 1, "roles": [...] }`. A new run records the selected preset and
@@ -47,11 +47,12 @@ Role IDs:
47
47
  only gate-authorized health escalations.
48
48
  - `axstack-debug-investigator-1..4` each probe one L1 brief.
49
49
 
50
- Provenance is matched on provider/model ID; record effort but never use it to
51
- create a mapping. Provenance absent from the preset's table row is
52
- unsupported and `INCOMPLETE`; report the exact gap and ask the user. Never
53
- derive a reverse pairing from slot position, driver, owner, or provider.
54
- Author and owner never review their own work.
50
+ Provenance is matched on provider/model ID; effort never maps. Missing table-row
51
+ provenance is unsupported and `INCOMPLETE`; report it and ask the user. Never
52
+ infer from slot, driver, owner, or provider. Author and owner never review.
53
+
54
+ The `axstack-implement` loop requires `mixed`; single-provider presets hold at
55
+ step (3) for user routing, with no substitution or same-provider review.
55
56
 
56
57
  ## Direct routes (no spec ceremony)
57
58
 
@@ -76,12 +77,11 @@ Author and owner never review their own work.
76
77
  handoff guide, and require explicit recipient acceptance before ownership
77
78
  changes. Missing capability is a setup gap; never invent one.
78
79
  - Colleague PR review -> `axstack-review`, peer mode.
79
- - Own PR maintenance or monitoring -> `axstack-review` in authored mode,
80
- `axstack-watch` for adoption.
81
-
82
- Research, explanation, improvement discovery, debugging, handoff, peer review,
83
- and adopted maintenance need no alignment, spec, or ticket map; authority and
84
- intent boundaries still apply.
80
+ - A status question about an own open PR or stack ("check now", "what's left",
81
+ "are we done", or "is it approved") -> `axstack-watch` in observation-only
82
+ mode. Explicit "address", "patch", or "fix" grants authorized maintenance.
83
+ - Other own PR work -> `axstack-review` authored mode or `axstack-watch`
84
+ adoption.
85
85
 
86
86
  ## Proportional scope identity
87
87
 
@@ -126,5 +126,4 @@ not alone a formal spec trigger. Hold affected unsafe work while reassessing.
126
126
  author provenance — then use `axstack-review` and `axstack-watch` without
127
127
  repeated approval or new spec ceremony. Never infer the author from the
128
128
  orchestrator or assume an imported own PR's author.
129
- - Direct later phase: start there and pass that phase's identity check; entry
130
- never admits work a deeper phase rejects.
129
+ - Direct later phase: start there and pass that phase's identity check.
@@ -161,5 +161,5 @@ approval; record chosen document names and paths once per run.
161
161
  actual dispatch. Alignment completion never dispatches a recipient.
162
162
 
163
163
  Alignment stops for both sizes only when the handoff is usable, its next scope
164
- identity is explicit, and execution has not started. The user invokes `axstack`
165
- to execute.
164
+ identity is explicit, and execution has not started. The user invokes
165
+ `axstack-implement` to execute.
@@ -5,11 +5,10 @@ description: When an approved task is ready to build or repair, use axstack-impl
5
5
 
6
6
  # Implement
7
7
 
8
- Deliver one reviewable candidate at an exact revision. Normal behavior changes
9
- have real red -> green -> refactor evidence; a narrowly accepted
10
- structure-preserving change has old-green characterization evidence. Name
11
- unverified boundaries and keep ownership unambiguous. Review and merge are
12
- later phases.
8
+ From an accepted scope identity, drive its task/PR map through author -> review
9
+ -> repair until every required PR is merge-ready or held. Keep exact revisions,
10
+ strict TDD evidence, ownership, and unverified boundaries explicit. The human
11
+ merges; the same run later reconciles those merges and closes out.
13
12
 
14
13
  ## 1. Admit the work
15
14
 
@@ -34,6 +33,9 @@ Independently confirm the applicable
34
33
  failure) returns to the author and does not increment the bug's fix ledger.
35
34
  - An adopted own-PR repair has its accepted maintenance snapshot.
36
35
 
36
+ If substantial work lacks an approved spec or matching ticket map, report that
37
+ exact gap, name `axstack-align` as the next route, and stop.
38
+
37
39
  Pin the exact base and current candidate revision. A missing, mismatched, or
38
40
  materially changed but unaccepted identity holds affected work; safe
39
41
  investigation may continue under the standing contracts. Proceed only with a
@@ -42,11 +44,9 @@ valid recorded identity and revisions; otherwise report the hold and exact gap.
42
44
  ## 2. Establish one owner and one writer
43
45
 
44
46
  For substantive delegated or resumable work, use the shared
45
- [run record](../axstack/references/run-record.md). On restart, reconcile it
46
- against actual Orca Tasks, Dispatches, sessions, Git revisions, GitHub state, the approved scope,
47
- Linear issue state, and watch registrations. Reuse the existing owner and
48
- author when valid. Ambiguous launch state is a hold on creating another writer,
49
- not evidence that the old writer disappeared.
47
+ [run record](../axstack/references/run-record.md). Reconcile it on restart with
48
+ the approved scope, Orca and forge state, exact revisions, tickets, and watches.
49
+ Reuse valid owners and authors; ambiguity holds a replacement writer.
50
50
 
51
51
  At execution start, bind work to the driver-owned Orca Run and one authoritative
52
52
  Task/Dispatch attempt. Preserve the actual IDs and process completion deliveries
@@ -57,30 +57,22 @@ Immediately before an actual role dispatch, read and follow the
57
57
  [Orca runtime boundary](../axstack/references/orca-runtime.md). Ordinary local
58
58
  reading and writing does not require that launch reference.
59
59
 
60
- One persistent owner remains accountable for the PR, fixes, evidence, and
61
- monitoring. Exactly one author writes a candidate at a time; accepted review
62
- repairs return to that author when its evidence is still usable. When the owner
63
- delegates writing, the owner does not edit that candidate concurrently. An
64
- ownership transfer occurs only when explicitly requested; follow the shared
65
- lifecycle's native capability preflight for that transfer. An ordinary restart
66
- or resume reconciles the existing sessions and run record without creating a
67
- fresh recipient.
68
-
69
- There is no fixed active-PR count. Fanout is dependency- and capacity-driven
70
- within configured host resource and spending limits, while one host owns the
71
- run and one writer owns each candidate. The driver queues conflicting or
72
- dependent work and coordinates dependent PRs through `gh stack`. Routine shape,
73
- split, fanout, and exception choices are autonomous driver decisions within the
74
- approved spec; size alone never requires user approval. A dependent candidate
75
- starts from its reviewed parent. When a reviewed parent changes, hold reliance
76
- on stale child evidence and child merge readiness. Rebase the child onto the
77
- new parent revision, re-run affected checks, and remeasure shape against the new
78
- actual base. Re-record the shape and re-check its level-matching rationale; size
79
- growth alone is not an automatic hold. A green parent does not prove the
80
- combined stack, but the parent need not wait for an independently reviewed
81
- child. Dispatch only when ownership, worktree, dependency revisions, writer
82
- exclusivity, and configured capacity agree with live state. Escalation occurs
83
- only if a split exposes an existing shared-contract hold.
60
+ One persistent owner remains accountable for each PR. Exactly one author writes
61
+ it; accepted repairs return there while its evidence is usable, and the owner
62
+ never edits concurrently. Only an explicit accepted transfer changes ownership;
63
+ ordinary resume reconciles the same sessions and record.
64
+
65
+ Fanout follows dependencies and capacity within configured limits. Queue
66
+ conflicts and dependent work; use `gh stack`, starting each child from its
67
+ reviewed parent. When the reviewed parent changes, hold reliance on
68
+ stale child evidence and child merge readiness; rebase onto the new parent revision,
69
+ re-run affected checks, and remeasure shape against it.
70
+ Size growth alone is not an automatic hold.
71
+ A parent need not wait for an
72
+ independently reviewed child. Dispatch only when ownership, worktree, dependency
73
+ revisions, writer exclusivity, and capacity agree with live state. Shape, split,
74
+ fanout, and exceptions are autonomous driver decisions within the approved scope.
75
+ Size alone never requires user approval.
84
76
 
85
77
  ## 3. Establish test-first evidence
86
78
 
@@ -95,11 +87,6 @@ restating source text or mirroring the intended implementation. Execute the
95
87
  check before changing production behavior and capture the expected behavioral
96
88
  failure. A missing-module error or unrelated setup failure is not red.
97
89
 
98
- For example, retry the same payment ID and observe one charge through the
99
- public interface. Counting internal helper calls alone would not prove that
100
- behavior. This illustrates the boundary test; it does not require a payment
101
- scenario in unrelated work.
102
-
103
90
  If no meaningful test-first check can be established, report why and hold
104
91
  dependent implementation for a scoped decision. Historical tests added after
105
92
  code remain noncompliant; they never become retroactive TDD evidence.
@@ -142,7 +129,7 @@ evidence when relevant. Name every unavailable OS, harness, credential, or
142
129
  other boundary instead of implying coverage.
143
130
 
144
131
  After the last change, pin the exact candidate revision and return this compact
145
- implementation receipt to the owner or driver:
132
+ implementation receipt to the driver:
146
133
 
147
134
  ```text
148
135
  Record: <progress.md path or tiny-task brief>
@@ -158,7 +145,55 @@ Next: <owner reconciles receipt, uses gh stack to push exact revision, confirms
158
145
  remote readback, then routes it to axstack-review>
159
146
  ```
160
147
 
161
- The author stops at that receipt and does not push. The owner follows the
162
- candidate-publication boundary without editing the candidate, and review starts
163
- only after remote readback confirms the exact revision. This grants no merge
164
- authority; the human merges by default.
148
+ The author stops at that receipt and does not push. The driver reconciles it,
149
+ uses `gh stack` to publish, confirms remote readback, and continues the loop
150
+ without editing the candidate. No step grants merge authority.
151
+
152
+ ## 6. Loop until merge-ready
153
+
154
+ Inputs are one snapshotted small-change intent or an approved spec and ticket
155
+ map. The unit is that accepted task/PR map: run independent PRs in parallel
156
+ within the fanout rule; run a dependent `gh stack` bottom-up, each child from
157
+ its reviewed parent. The driver is owner, sole record writer, dispatcher,
158
+ publisher, and wait-holder for every loop PR it creates. Resume preserves an
159
+ existing live owner absent an accepted transfer.
160
+
161
+ For each PR:
162
+
163
+ 1. Dispatch `axstack-author` under §§3-5 and consume its strict-TDD receipt.
164
+ 2. Publish through candidate-publication and read back the exact SHA.
165
+ 3. Dispatch and consume the authored-mode `axstack-review` selected from actual
166
+ author provenance.
167
+ 4. Route the verdict. `APPROVE` at that head plus `axstack-watch` §5's full
168
+ predicate—required checks, all feedback, approvals, mergeability, and
169
+ exact-revision receipts—records `merge-ready`. With required checks pending,
170
+ use the forge-native blocking check wait, bounded and used once per revision, then
171
+ re-evaluate. Timeout, error, or missing wait capability records `held` at
172
+ that revision with reason and resume condition; it never triggers author
173
+ repair. Notify “checks pending, resume when green”, not “decision needed”.
174
+ `REQUEST_CHANGES`, a failed required check, or post-readiness feedback returns
175
+ findings to the same author for a new revision, increments `repairs`, and
176
+ returns to step 1. `INCOMPLETE`, a provenance gap, unavailable model, serious
177
+ risk, or the third `REQUEST_CHANGES` on one PR records `held`. A changed
178
+ parent sends its child back to step 1.
179
+
180
+ One run-level completion wait covers every unsettled Dispatch; the bounded
181
+ forge check wait is the only other wait. End a turn only when every required PR
182
+ is `merge-ready` or `held`, after notification (b) or (a). Raise serious risk
183
+ (c) immediately when found. Notifications use `axstack-relay` under the recorded
184
+ Notification policy: (a) a user-decision hold, (b) the merge-ready set and the
185
+ all-merged event—two per run—and (c) serious risk; never progress.
186
+
187
+ Merge-ready is the human boundary: the user merges, bottom-up for a stack. The
188
+ driver resumes on the user's next message or `/axstack-watch`; no Orca merge
189
+ wake exists today. Re-read forge state: record forge-merged PRs as `merged`;
190
+ changed heads or feedback return to step 1; release nothing before Close-out.
191
+ Run Close-out once only after every required PR is forge-merged and acceptance
192
+ passes. It settles workers, records counts, makes the auditor decision and
193
+ settlement, releases worktrees, closes eligible tickets, and archives the run.
194
+
195
+ The loop requires the `mixed` two-provider authored-review row. `codex-only` or
196
+ `claude-only` holds at step (3) for an explicit user routing choice, with no
197
+ substitution or same-provider review. Derived PR states are `authoring |
198
+ published | in-review | repairing(n) | merge-ready | merged | held`. The run is
199
+ done only when every required PR is forge-merged and Close-out has receipts.
@@ -69,7 +69,8 @@ model-free and read-only, has no gate, and records `watchdog.log`; there is no
69
69
  watch deadline for automations. The driver is the automation session itself,
70
70
  with no `axstack-monitor` or `axstack-owner` role row; `axstack-monitor` stays
71
71
  an optional read-only observer that never sends. One read-only PR observation
72
- needs neither.
72
+ needs neither. The publishing driver is the live owner for a status check;
73
+ materialize no `axstack-owner` and start no automation for a read-only check.
73
74
 
74
75
  For standalone adoption, materialize `axstack-owner` only when no live owner
75
76
  exists. Once it exists, the current chat is not a competing coordinator. Only
@@ -157,3 +158,6 @@ Resume: <known commands or verified refs needed to reconcile from this revision>
157
158
 
158
159
  The watch ends only when registrations are stopped, receipts are recorded, and
159
160
  the PR is either merged or represented by this resumable state.
161
+ When every required PR is merged, follow the lifecycle
162
+ [Close-out](../axstack/references/lifecycle.md#close-out) before reporting the
163
+ run as done.
package/src/installer.js CHANGED
@@ -247,12 +247,14 @@ export async function validateBundle(bundleDir, selectedPreset = null) {
247
247
 
248
248
  const files = [];
249
249
  for (const dir of skillDirs) {
250
- const skillMark = join(skillsRoot, dir.name, 'SKILL.md');
251
- try {
252
- const s = await stat(skillMark);
253
- if (!s.isFile()) throw new Error();
254
- } catch {
255
- throw new Error(`skill ${dir.name} is missing SKILL.md`);
250
+ if (dir.name !== 'axstack') {
251
+ const skillMark = join(skillsRoot, dir.name, 'SKILL.md');
252
+ try {
253
+ const s = await stat(skillMark);
254
+ if (!s.isFile()) throw new Error();
255
+ } catch {
256
+ throw new Error(`skill ${dir.name} is missing SKILL.md`);
257
+ }
256
258
  }
257
259
  await walkSkills(join(skillsRoot, dir.name), skillsRoot, files);
258
260
  }
@@ -508,7 +510,7 @@ export async function installBundle({
508
510
  }
509
511
  instructionPlan = planInstruction({
510
512
  text: existingInstructionsRaw,
511
- block: renderInstructionBlock(skillsRoot),
513
+ block: renderInstructionBlock(),
512
514
  ownership: boundInstructions.path === instructionsFile ? boundInstructions : null,
513
515
  force,
514
516
  });
@@ -1,14 +1,14 @@
1
1
  // Pure planning and byte-preserving edits for the Axstack-owned routing block.
2
2
  import { hashContent } from './manifest.js';
3
- import { join } from './posixpath.js';
4
3
 
5
4
  const BEGIN = '<!-- axstack:begin v1 -->';
6
5
  const END = '<!-- axstack:end -->';
7
6
 
8
- export function renderInstructionBlock(skillsDir) {
7
+ export function renderInstructionBlock() {
9
8
  return [
10
9
  BEGIN,
11
- `Use Axstack for engineering work. Load \`${join(skillsDir, 'axstack', 'SKILL.md')}\` to route the request.`,
10
+ 'Use Axstack for engineering work: invoke the matching `axstack-*` skill directly.',
11
+ '`axstack-implement` loops author -> review -> repair until every PR is merge-ready.',
12
12
  'Route every subagent, delegated worker, reviewer, and cross-harness dispatch through Orca orchestration via the `orca` CLI and its `orca-cli` / `orchestration` skills so the work stays visible.',
13
13
  'Do not use a harness native subagent tool for delegated work.',
14
14
  END,
@@ -1,81 +0,0 @@
1
- ---
2
- name: axstack
3
- description: When routing an engineering run through Axstack, use axstack to select the applicable phase and scope identity.
4
- ---
5
-
6
- # Axstack entry
7
-
8
- Route the current request to one Axstack phase with the right scope identity.
9
- The current chat remains the driver; Orca owns runtime orchestration.
10
-
11
- For an explicit relay message or transport test, use
12
- [axstack-relay](../axstack-relay/SKILL.md) directly. No engineering scope
13
- identity or decision workflow is needed for that send. The same skill handles
14
- urgent or blocking notifications under an explicit standing instruction.
15
-
16
- ## Route the request
17
-
18
- 1. Classify the request with [Shared routing](references/routing.md). Direct
19
- research, explanation, improvement discovery, peer-review, adopted-watch,
20
- and handoff routes need no spec
21
- ceremony. Only an explicit user-requested ownership transfer can use the
22
- capability-gated native route in
23
- [Lifecycle and receipts](references/lifecycle.md#native-handoff-and-resume),
24
- not an Axstack handoff phase. Preparation completion, watch expiry, and
25
- ordinary resume update or reconcile the run record without launching it.
26
- 2. For new engineering work, validate scope identity before invoking any phase.
27
- Record `small`, `substantial`, or `unclear` plus a brief reason, then apply the
28
- [proportional scope identity](references/routing.md#proportional-scope-identity).
29
- A small clear change proceeds from its snapshotted small-change intent.
30
- Substantial work proceeds only from an approved spec and matching ticket
31
- map. Clarify unclear size before dispatch.
32
- 3. Only after validation passes, invoke exactly the selected phase. A directly
33
- invoked later phase starts there and must pass its own identity check. When
34
- substantial work lacks an approved spec or matching ticket map, return that
35
- exact gap, name `axstack-align` as the next route, and stop the current
36
- invocation; do not invoke align, spec, or tickets. Apply the same stop to a
37
- mismatched or invalidated identity. Never admit work that a deeper phase
38
- would reject.
39
-
40
- The route is settled when one applicable phase is named with its valid scope
41
- identity, or the exact preparation/setup gap is reported with affected work
42
- held.
43
-
44
- ## Load at the action boundary
45
-
46
- - Every independently called phase loads [Standing contracts](references/contracts.md),
47
- which requires lifecycle and audit loading before action.
48
- - Before an actual Axstack role dispatch, delivery, settlement, or handoff, load
49
- [Orca runtime](references/orca-runtime.md). Ordinary reading, writing, and
50
- local checks do not require launch discovery.
51
- - Substantive delegated or resumable work uses the
52
- [Local run record](references/run-record.md).
53
- - When the current session is an Orca PR automation (driver or watchdog), load
54
- [Automation sessions](references/automations.md) before any discovery,
55
- review, gate, or mutation.
56
- - When review escalation or watch notification is eligible and the brief has a
57
- `Notification policy`, use the optional
58
- [axstack-relay](../axstack-relay/SKILL.md); otherwise keep notification in
59
- the current Orca conversation.
60
-
61
- ## Lifecycle
62
-
63
- This is a phase map, not an automatic dispatch sequence.
64
-
65
- 1. `axstack-align` settles substantial scope and decisions.
66
- 2. `axstack-spec` creates the single user-approved execution baseline.
67
- 3. `axstack-tickets` maps capabilities, tasks, and dependencies, then
68
- preparation stops with a resumable handoff.
69
- 4. `axstack-implement` produces owned candidates with strict TDD.
70
- 5. `axstack-review` gives peer PRs the two configured independent same-brief
71
- reviewer roles; authored PRs get one complete eligible non-author/non-owner
72
- review based on actual author provenance and the routing snapshot.
73
- 6. `axstack-watch` monitors within the shared deadline and hands off remaining
74
- work.
75
- 7. The human merges by default, bottom-up for a stack. Review approval never
76
- grants merge authority.
77
-
78
- Autonomous progress, model holds, serious-risk handling, mutation authority,
79
- and the one-host ownership contract live in
80
- [Standing contracts](references/contracts.md). Load only the selected phase
81
- and the references its action requires.