axstack 0.19.0 → 0.20.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,474 +1,167 @@
1
- # Automation sessions
2
-
3
- Read this when the current session is the native Orca PR **driver** or its
4
- **watchdog**. The approved contract is
5
- `docs/specs/pr-automations.md` revision 6; this reference restates the parts an
6
- automation session must execute and does not widen them.
7
-
8
- Pair A/B is retired for this contract; its artefacts remain untouched.
9
-
10
- ## Roles and authority
11
-
12
- - **driver** — every 15 minutes, fresh Opus session in the host's `root`
13
- folder workspace, which is not a git repository and belongs to no project. It discovers GitHub work, binds the persistent Orca
14
- Run, creates one worktree per selected PR and checks its head out there, dispatches the
15
- matching Axstack agent, reconciles completions and decisions, then exits. The
16
- driver is the automation session itself, with no `axstack-monitor` or
17
- `axstack-owner` role row. It never performs review or repair in its own
18
- session. Its only direct GitHub mutation is an already-approved decision:
19
- one `gh pr review` or one fast-forward push of a preserved candidate.
20
- - **watchdog** — hourly at a distinct minute. It is a shell precheck with no
21
- model, never launches an agent session, never mutates GitHub, and only sends
22
- the health notifications defined below.
23
- - `axstack-monitor` remains an optional read-only observer that never sends.
24
-
25
- Self is resolved on every run with `gh api user --jq .login`; never hardcode
26
- it. The review allowlist is `defi-com/monorepo`, `defi-com/mobile`,
27
- `defi-com/azure-next-hybrid`, and `defi-com/ci-workflows`. The repair allowlist
28
- is `defi-com/monorepo` and `defi-com/mobile`. State both lists verbatim in the driver prompt. An own PR is
29
- open, authored by self, on the repair allowlist, and not a draft. A peer PR is
30
- open, on the review allowlist, authored by someone else, and either officially
31
- review-requested from self or has a non-self comment that both mentions self
32
- and asks for review or response. The driver reads and records the comment id
33
- and its interpretation; incidental mentions are discovery only.
34
-
35
- Outside the union of both allowlists, record only and exclude the PR from the
36
- fingerprint. Peer code is read-only, though its worktree may install
37
- dependencies and run repository tests. Own-PR authority permits repair, test,
38
- commit, and fast-forward `git push` without lease or force. Force-push, rebase,
39
- merge, close, every `gh stack` sync/restack/rebase/merge/link/submit action,
40
- `COMMENT` reviews, and every GitHub write by the watchdog or Hermes are
41
- prohibited.
42
-
43
- ## Discovery and wake
44
-
45
- The bounded driver precheck contains no model and exits 0 only for changed or
46
- due work.
47
-
48
- 1. If `cursor.json.tick_started_at` is newer than `tick_done_at` and younger
49
- than 1 h, append `running` to `precheck.log` and exit 3. This overlap guard
50
- inspects no terminal.
51
- 2. Resolve self, then run exactly four `gh search prs --state open --limit 100
52
- --json url,number,repository,updatedAt` searches: `--author @me`,
53
- `--review-requested @me`, `--mentions @me`, and
54
- `--reviewed-by @me --review changes_requested`. A result count equal to 100
55
- is truncation: append `error`, exit 2, and do not write a fingerprint.
56
- Deduplicate by URL with own-PR precedence and drop repositories outside the
57
- allowlist union.
58
- 3. For each retained PR, read `gh pr view --json headRefOid,baseRefName,isDraft,
59
- statusCheckRollup,author,latestReviews` and resolve the base SHA with
60
- `gh api repos/<repo>/commits/<base>`. Own PRs contribute head, base, draft,
61
- checks reduced to `{name, conclusion|state}` pairs, and the set of
62
- `latestReviews` `{id, state}` whose `commit` is the head; others contribute
63
- head and base.
64
- 4. Debounce peer and fourth-search heads: they enter the hashed subset only on
65
- the second consecutive precheck that observes them. The first observation
66
- goes in `pending.json.seen[]`. Own heads enter immediately. Hash this subset
67
- as the fingerprint.
68
-
69
- ### Due control work
70
-
71
- Read `cursor.json` and `decisions/`. Work is due for:
72
-
73
- - a decision in `approved` or `rejected` without `consumed_at`;
74
- - a `spent` decision without a receipt, which needs reconciliation;
75
- - a `deferred[]` entry whose head still matches discovery;
76
- - an expired repair cap;
77
- - a dispatch marker older than 3 h;
78
- - an unsettled Orca delivery in `pending_settlement[]`;
79
- - a `runtime_refusal` record, because its re-test needs a launched tick.
80
-
81
- An `open` decision is not due. Write `pending.json` with the fingerprint,
82
- `observed_at`, `seen[]`, full discovery list, and hashed subset. Append
83
- `<ts> changed|due|unchanged|running|error` to `precheck.log`; exit 0 for
84
- `changed` or `due`, 1 for `unchanged`, 2 for `error`, and 3 for `running`.
85
-
86
- ## Driver tick
87
-
88
- The driver performs this order and exits:
89
-
90
- 1. If `tick_started_at` is newer than `tick_done_at`, add `previous tick did
91
- not finish` to `cursor.json.health[]`. Write `tick_started_at`. Match this
92
- session's `ORCA_TERMINAL_HANDLE` to the driver's `terminalPtyId` from
93
- `orca automations runs --id <driver>`, then verify the transcript model.
94
- Non-Opus or unknown identity records a hold and `tick_outcome: held`, and
95
- dispatches nothing.
96
- 2. Bind the persistent Run with `orca orchestration run-use` and read the
97
- inbox. For a `worker_done` matching a live marker, verify the review id at
98
- the bound head, push range, or opened token. Release the worker; once its
99
- release receipt is settled and the process has exited, run the cleanup
100
- under "Run directory" below, which proves the worktree disposable before
101
- resetting it; then clear the marker. Unverifiable delivery stays in
102
- `pending_settlement[]` and blocks only that PR.
103
- 3. Consume decisions as their sole consumer under "Decision tokens" below.
104
- 4. Reconcile every marker older than 3 h. A live worker gets `worker-stop`; an
105
- exited worker gets `worker-abandon`. Unknown liveness or user takeover
106
- retains worktree and marker and blocks only that PR. An abandon is
107
- confirmed by its accepted abandon receipt plus proven process exit; there
108
- is no release receipt on this path. After confirmed abandon, run the same
109
- cleanup under "Run directory" below with its extra clean-tree condition — a
110
- dirty or unproven worktree is retained, not removed — append one health
111
- line, increment the head's `abandon_count`, and drop that head so normal
112
- selection retries once. For a
113
- review-triggered dispatch, use the marker's `trigger` to remove exactly its
114
- review id and digest from `processed_reviews[]`. A second abandon at that
115
- head is a user-owned hold.
116
- 5. Select changed PRs and matching `deferred[]` entries oldest `updatedAt`
117
- first. Skip a current-head decision in `open` or `approved`, and skip a live
118
- marker. Dispatch within the repair and review rules below.
119
- 6. Promote the `pending.json` fingerprint verbatim, because it records
120
- observation rather than completion. Write `tick_done_at` and
121
- `tick_outcome: ok`, then close your own terminal tab and exit: each tick is
122
- a fresh session and there is no hygiene sweep, so a tab left open outlives
123
- the tick as a stray terminal in the root workspace.
124
-
125
- ## Dispatch and repair selection
126
-
127
- Claude Code trusts a folder per git toplevel and stops at its "Quick safety
128
- check" dialog otherwise, and the driver never answers that dialog for a
129
- worker. Trust inherits from the project's primary clone, which the user
130
- trusted once (a fresh child worktree launched with no dialog — canary
131
- 2026-09-18). There is no fixed pool: fetch the head into the project clone first (a
132
- failed fetch is a health line and no dispatch; no worktree exists yet), then
133
- create one Orca worktree per dispatch, `orca worktree create --repo id:<clone
134
- id> --name <repo short>-<num>-<head7> --base-branch <default branch>
135
- --parent-worktree id:<clone id>::<clone path> --setup skip` (a name collision
136
- with a retained worktree at the same head appends the tick's
137
- `tick_started_at` stamp); close the creation terminal Orca opens in it with
138
- `orca terminal close --worktree <selector> --all`; check the worktree out
139
- detached at the pinned head and verify HEAD equals it; the marker's worktree
140
- is that path. Never write
141
- `~/.claude.json`. The worker is launched by Orca itself — `worker-start
142
- --agent claude --model claude-opus-5 --effort medium` in that worktree — so
143
- the runtime owns the process: `worker-release` ends it and `worker-show`
144
- proves it exited, which is what allows the worktree to be removed. Never
145
- pre-create the worker's terminal or hand a terminal handle to
146
- `worker-start`: a reused handle is a resource Orca labels `external`, one it
147
- can neither stop nor prove exited, so every such worktree ends retained.
148
- After `worker-start` run `worker-show` on the receipt's dispatch id and
149
- require `projection.resource.state == owned`; anything else is a launch Orca
150
- does not own: apply the runtime-refusal recovery rules under "Safety holds"
151
- (the `worker-list` row's `nextAction` argv verbatim; `none` means inspect
152
- and retain), append a `worker not owned` health line and a `retained_slots[]`
153
- entry for the worktree, defer the PR, and record no marker. Orca's per-agent
154
- default arguments supply `--dangerously-skip-permissions`; the brief loads
155
- the skill files it needs by path. If `worker-start` reports a failed stage
156
- or a visible hold (the "Quick safety check" trust dialog) the worktree is not
157
- trusted: name it in a health line, defer the PR, dispatch nothing, never
158
- answer the dialog, and remove the worktree through the no-worker branch.
159
- Project customizations load as they would for the user; the allowlist is
160
- defi-com only and the user accepted that surface on 2026-09-18.
161
-
162
- Every selected PR receives one dispatch marker with task id, dispatch id,
163
- worktree, head, `started_at`, reservation (`verdict` or `repair`), and trigger:
164
- `{kind: check, name, app_id}` or `{kind: review, review_id, digest}`. There is at
165
- most one live marker per PR. The marker is the claim shared by scheduled and
166
- attended sessions; revalidation immediately before an external call is its
167
- second half. It makes no exactly-once claim against concurrent human GitHub
168
- activity.
169
-
170
- An own PR needs repair when either trigger applies:
171
-
172
- 1. a failing check has a base check-run with the same `name` and the same
173
- producing `app.id` observed passing through
174
- `gh api repos/<repo>/commits/<base>/check-runs`; for a legacy commit status,
175
- its counterpart has the same `context`. A missing, pending, or same-name
176
- different-app base check holds repair;
177
- 2. a `CHANGES_REQUESTED` review at the current head, by any account, has both
178
- a review id absent from `cursor.json.processed_reviews[]` and a SHA-256 body
179
- digest not recorded for that PR and head. Both keys are required: the same
180
- finding under a new review id must not re-trigger repair. Record review id
181
- and body digest when dispatching. A superseded head with a new review
182
- triggers again at the new head.
183
-
184
- Repair also requires no deploy-on-push head branch, a head not already in
185
- `repaired_heads[]`, and selection of the lowest own PR in its stack that
186
- needs repair. Create the
187
- dispatch worktree (parented to that project's primary worktree, so the work
188
- appears under the project it serves) at the exact head, and dispatch one
189
- `axstack-watch` agent in authored repair mode. Its
190
- brief contains only the triggering checks or review findings. Each open
191
- descendant records one user-owned `pending restack` hold until it stops needing
192
- repair. At dispatch record the head in `repaired_heads[]` (`pr`, `head`,
193
- `dispatched_at`): one repair per head, no time cap. A repair pushes a new
194
- head; a still-not-merge-ready new head shows a new failing check or review
195
- and is repaired again; a repair that pushes nothing is not retried at that
196
- head until a human or a new commit moves it. A confirmed abandon removes the
197
- record (retry once).
198
-
199
- A debounced peer PR is eligible when self has not reviewed its head. Read
200
- `gh pr view --json reviews` before dispatch. Whenever any self review with
201
- state `CHANGES_REQUESTED` exists on the PR, from whichever search it came,
202
- dispatch only if the latest effective, non-dismissed self review body contains
203
- the line prefix `<!-- axstack-automation verdict` or its id is listed in
204
- `cursor.json.legacy_automation_reviews[]`; otherwise record and skip, so a
205
- human-placed block is never overwritten. A dismissed block and a self-approved
206
- PR are skipped. Dispatch one `axstack-review` agent in peer mode and link the
207
- prior review in its brief.
208
-
209
- The only concurrency limit is the host-wide cap: at most eight live dispatch
210
- markers across all repositories and both reservations, oldest eligible
211
- first; plus one repair per head. Put every eligible PR not dispatched
212
- because the cap is reached in `deferred[]` with repo, PR, and head. Reaching
213
- the cap records the count and is not a hold.
214
-
215
- ## Agents and verdicts
216
-
217
- Every agent works in its own project-local worktree checked out detached at
218
- the exact head and reports only through the Orca worker protocol.
219
-
220
- Peer review runs the two isolated configured reviewers on the identical brief,
221
- then the Luna gate. `APPROVE` requires complete exact-head/base reviews, gate
222
- `proceed`, and zero validated blockers. `REQUEST_CHANGES` requires the same
223
- completeness and `proceed`, plus at least one evidenced blocking finding.
224
- `INCOMPLETE`, unresolved disagreement, unavailable review or gate, or unknown
225
- GitHub state publishes nothing. Immediately before `gh pr review`, the agent
226
- re-reads self's reviews at the head and skips with the existing id when one is
227
- already present, then re-checks head, base, draft, authorship, and allowlist.
228
- The verdict body ends with this exact marker line:
229
-
230
- ```text
231
- <!-- axstack-automation verdict head=<sha> -->
232
- ```
233
-
234
- The verdict body is written for the person reading the PR and reads as one
235
- reviewer's findings. It never names the reviewer count, the brief, the
236
- angles, the gate, receipts, or which reviewer found what: "two independent
237
- reviews", "from secondary", and "the reviewers ran" are pipeline facts, not
238
- review content. It is self-contained: every validated finding, blocking or
239
- not, appears in full in the body — evidence and consequence, file and line
240
- where they exist — so a reader is never told that notes exist without seeing
241
- them, and a finding is never dropped to keep the body short. It never points
242
- at the local review file or at anything the reader cannot open. Evidence
243
- appears as what was checked and observed, not as who ran it: the reviewed
244
- head and base SHAs, CI status, test counts, and diff size are reader-useful
245
- facts and belong; "shape verified" and "pinned CI" are pipeline phrasing and
246
- do not. The marker line is the only pipeline artefact the body carries.
247
-
248
- Write the local review file to the workspace review directory
249
- `~/defi/misc/reviews/` under the existing convention:
250
- `review-PR-<num>.html` with no prefix means `defi-com/monorepo`;
251
- `review-mobile-PR-<num>.html`,
252
- `review-azure-next-hybrid-PR-<num>.html` and `review-ci-workflows-PR-<num>.html`
253
- name the other repositories. A write
254
- failure is recorded but does not withhold the verdict.
255
-
256
- Authored repair commits a local candidate, obtains one Sol review at that
257
- local SHA and the Luna gate, resolves every validated blocker, records test
258
- evidence, re-reads remote head/base/draft/deploy set/allowlist, then pushes
259
- fast-forward. The monorepo unit gate uses `~/.bun-1.2.2/bin/bun`; a suite that
260
- cannot run locally is an explicit unverified boundary.
261
-
262
- Every reviewer brief ends exactly:
263
-
264
- ```text
265
- Escalate to user: yes | no — <criterion> — <reason>
266
- ```
267
-
268
- PR work has exactly three criteria: a security concern, a permanent on-chain
269
- state change, or an architectural change in approach. There is no
270
- automation-health criterion for reviewers. The Luna gate returns exactly one
271
- token, `escalate` or `proceed`; there is no gate for health findings.
272
- `escalate` opens a decision token, sends its message, and exits without waiting
273
- for a reply. `proceed` does not override a validated blocker. `worker_done`
274
- names the PR, head, action, GitHub receipt or opened token. The driver alone
275
- writes the run record.
276
-
277
- ## Decision tokens
278
-
279
- Store `decisions/<token>.json`, where `<token>` contains at least 96 random
280
- bits as lowercase hex, produced for example by `openssl rand -hex 16`.
281
- Immutable bound fields are written once: `repo`, `pr`, `head`, `base`,
282
- `action` (`approve-verdict`, `request-changes-verdict`, or `push`), plus exact
283
- `body` and `commit` for verdicts or `candidate_sha`, `head_branch`, and
284
- `expected_remote_head` for a push. Mutable fields are `state`, `created_at`,
285
- `decided_at`, `decided_message_id`, `consumed_at`, `receipt`, `send`, and
286
- `reason`. Never delete a token file.
287
-
288
- Lifecycle has one named writer per transition, each by temp file + rename:
289
-
290
- | transition | writer |
291
- | --- | --- |
292
- | create `open` | the agent that escalated |
293
- | `open → approved` / `open → rejected` | the Hermes script only |
294
- | `approved → spent` / `approved → stale` | the driver only |
295
- | `rejected → closed` | the driver only |
296
- | `spent` + `receipt` | the driver only |
297
-
298
- ### Opening
299
-
300
- Before opening a `push` token, the repair agent pins its candidate with local
301
- ref `refs/axstack/decisions/<token>` in the project clone, so the candidate
302
- survives worktree removal. It then sends one
303
- `hermes send --to telegram` message naming the PR, criterion, every reviewer's
304
- reason, and the exact replies `/axstack-decide approve <token>` and
305
- `/axstack-decide reject <token>` (the slash form loads the Hermes skill
306
- deterministically; bare `approve <token>` is best effort). Store the
307
- send receipt. The driver retries a `failed` send once next tick and reconciles
308
- an `uncertain` send against Hermes output before any retry.
309
-
310
- The Hermes gateway's fixed `axstack-decide` script accepts only those two exact
311
- commands. It validates the private user/channel configuration on every call,
312
- restricts tokens to `^[0-9a-f]{24,}$`, locks and re-reads an `open` token,
313
- writes only its decision fields by temp file and rename, and prints one line.
314
- Hermes never runs `gh`, `git`, or `orca`.
315
-
316
- ### Consuming
317
-
318
- The driver is the only consumer. Immediately before acting it re-reads that
319
- `state == approved`, then revalidates open/unmerged state, bound head/base,
320
- absence of a self review for a verdict, or expected remote head plus reachable
321
- candidate for a push. It writes `spent` before the GitHub call, executes
322
- exactly the bound action with no second gate, then stores the receipt. A spent
323
- token without a receipt is
324
- reconciliation: match the exact review commit/body or destination
325
- ref/candidate on GitHub; record a match, or the driver retries once under the
326
- same approval after proving non-execution, or hold ambiguity. A revalidation
327
- failure writes `stale`
328
- with the reason and drops the PR's head from `cursor.json`; the driver never
329
- mints a token. Close rejected files and skip that head. After `spent` or
330
- `stale`, delete the candidate ref. Tokens do not expire; the watchdog reports
331
- one open longer than 24 h.
332
-
333
- ## Watchdog
334
-
335
- The hourly shell precheck only reads `cursor.json`, `precheck.log`,
336
- `decisions/`, and `orca automations runs --id <driver>`. It performs exactly
337
- these four checks and always exits non-zero, so no model session launches:
338
-
339
- | check | trips when |
340
- | --- | --- |
341
- | driver stuck | a `changed` or `due` precheck is older than 1 h with no later `tick_done_at`, including a driver that never wrote `tick_started_at` (Orca run status alone is not evidence of completion) |
342
- | precheck failing | the last three `precheck.log` entries are `error` |
343
- | worker stuck | a dispatch marker is older than 3 h 30 min and remains uncleared |
344
- | decision waiting | an `open` decision is older than 24 h |
345
-
346
- Each trip is `(check, first_observed)`. Write one JSON line per tick to
347
- `watchdog.log` with `{ts, checks, trips, sent}` and send each occurrence once
348
- through `hermes send --to telegram`. Persist `sent`, `failed`, or `uncertain`;
349
- retry `failed` next tick and reconcile `uncertain` before retry. Re-send only
350
- after the check was observed clear and later recurs. Unreadable evidence is an
351
- `unknown` occurrence. A healthy tick writes its line and sends nothing. There
352
- is no gate for health findings.
353
-
354
- ## Run directory
355
-
356
- The driver and watchdog run from the host's `root` folder workspace, not a
357
- project worktree: no project owns the automation, and every dispatch
358
- worktree belongs to the project it serves. That workspace is not a git repository, so
359
- the run directory is private host state, one
360
- `~/.local/share/axstack/runs/<run id>/` directory. Settlement leaves nothing
361
- behind, but never destroys work. Before any destructive step the driver
362
- proves the worktree is disposable: the worker is settled — on the release path
363
- a settled release receipt, on the abandon path an accepted abandon receipt,
364
- either with proven process exit; pending or unknown stops here — the
365
- worktree's HEAD is either the pinned head or a candidate that is durably
366
- reachable — pushed to the head branch, tested only after a successful
367
- targeted fetch of that exact remote branch into a per-dispatch ref, never
368
- `FETCH_HEAD`, so neither a stale tracking ref nor a concurrent fetch in the
369
- shared clone can fake durability, a failed fetch retaining the worktree — or held by a
370
- `refs/axstack/decisions/<token>` ref in the project clone; and, on the
371
- abandon path, the worktree has no uncommitted changes.
372
- Only then it closes any terminal tab still listed and removes the worktree
373
- with `orca worktree rm --worktree <selector> --force` (the proof is the
374
- gate; a finished worker's untracked artefacts are not), so the dispatch
375
- leaves nothing behind. On the settled path the worker has finished, so untracked
376
- files are artefacts by definition; a candidate there is already pushed or
377
- token-held. A worktree whose release is settled but whose HEAD cannot be
378
- proven disposable, and an abandoned worktree that is dirty or holds an
379
- unproven candidate, are both **retained** — retained in place: named in one `health[]`
380
- line with path and SHA, and blocking only that PR with the user as owner.
381
- Retention is mechanical on both paths: the driver appends `retained_slots[]`
382
- `{slot, pr, head, reason}` (`slot` is the worktree path); a retained worktree
383
- is never removed by the driver, and the entry is cleared only by the user
384
- after reconciling the candidate, who also removes the worktree. A leftover
385
- worktree or terminal that outlives its dispatch without such a retention
386
- record is a health finding. The run directory contains:
387
-
388
- - `cursor.json` — driver only, with these exact keys: `fingerprint`,
389
- `tick_started_at`, `tick_done_at`, `tick_outcome`, `prs{url: {head, base,
390
- draft, checks, reviews, last_self_review}}`, `dispatch_markers[]` (`pr`,
391
- `task_id`, `dispatch_id`, `worktree`, `head`, `started_at`, `reservation`,
392
- `trigger`), `deferred[]`, `pending_settlement[]`,
393
- `retained_slots[]` (`slot`, `pr`, `head`, `reason`),
394
- `repaired_heads[]` (`pr`, `head`, `dispatched_at`),
395
- `repair_caps{url: {expires_at}}` (legacy: the first rev-6 tick clears it to
396
- `{}` with one health line; never written again),
397
- `abandon_count{head: n}`,
398
- `processed_reviews[]` (`review_id`, `pr`, `head`, `digest`),
399
- `deploy_on_push{repo: [branches]}`,
400
- `legacy_automation_reviews[]`, `health[]`,
401
- `runtime_refusal{code, first_seen, last_seen}` (absent when no runtime
402
- hold is open);
403
- - `pending.json`, `precheck.log` — driver precheck only;
404
- - `decisions/<token>.json` — writers assigned by the lifecycle table;
405
- - `watchdog.log` and `watchdog-state.json` (occurrence `first_observed` values
406
- and send receipts) — watchdog only;
407
- - `progress.md` — driver only, one line per tick plus holds and mention
408
- readings, with no per-PR prose.
409
-
410
- Timestamps are UTC `YYYY-MM-DDTHH:MM:SSZ`; an unparsable timestamp is an
411
- `error` for the precheck and `unknown` for the watchdog, never silently
412
- ignored.
413
-
414
- Orca run history is the authoritative log. The launch workspace is the host's
415
- `root` folder workspace, not a project worktree: nobody develops there, it is
416
- not a git repository, and no project owns the automation. The briefs and the
417
- escalation template are read from the axstack checkout at an absolute path
418
- given in the prompt, never relative to the launch workspace. Keep this notification-policy edge in the
419
- run record: `Notification policy` authorizes the token and watchdog sends;
420
- delivery uses [axstack-relay](../../axstack-relay/SKILL.md).
421
-
422
- ## Safety holds
423
-
424
- - Non-Opus or unknown driver identity records a hold and dispatches nothing.
425
- - A GitHub API error makes PR state unknown; never publish or push for it.
426
- - Enumerate deploy-on-push branches from both repair repositories before
427
- enabling and store them in `cursor.json`; re-check on allowlist changes.
428
- - A later tick observing resolution or an explicit user decision clears a
429
- hold. Silence never clears one. For a hold caused by an Orca runtime
430
- refusal — a sub-worker dispatch rejected for depth, a launch capability the
431
- runtime declines — observing resolution means re-attempting the refused
432
- operation, once per tick, on the next eligible PR: success clears the hold.
433
- The hold is keyed on Orca's structured error code (for the depth case,
434
- `nested_worker_depth_exceeded`), stored in `cursor.json` as
435
- `runtime_refusal {code, first_seen, last_seen}`; the same code keeps the
436
- hold and updates `last_seen` without a new health line, a different code is
437
- a new finding. Prose is never the key. When `worker-start` itself is
438
- refused after the worktree was created, there is no worker, so the
439
- settlement proof does not apply; the driver reads the receipt's `failedStage`
440
- and `residualResources` first. With no Dispatch and no residual resources
441
- the no-worker branch applies: the worktree's HEAD must equal the pinned
442
- head and `git status --porcelain` must be empty, and then the worktree is
443
- simply removed in the same tick.
444
- With a Dispatch or any residual resource the failed start owns runtime
445
- state, and retaining alone is not recovery: the driver follows the
446
- runtime's recovery guide. With a Dispatch: `worker-list` for that run, and
447
- the row's `nextAction` is an object `{kind, argv}` — a non-empty `argv` is
448
- run verbatim through the same Orca executable and nothing else, while
449
- `kind: none` authorizes no action beyond inspection and retention. With
450
- residual resources but no Dispatch there is no row: the mutation itself is
451
- recovered through `request-show` on the receipt's request id. The no-worker
452
- branch applies only after the resources are proven gone. It never retries
453
- in the same tick. Anything unproven retains the worktree with a health line
454
- naming the stage and the resources. Persisted configuration such
455
- as `orca-data.json` is never evidence either way; it is a snapshot that
456
- lags the live setting, and the driver never reads it.
457
-
458
- ## Cutover
459
-
460
- Perform this order: the new pair exists disabled; the amended skills and
461
- references are installed; C and D are disabled; every old driver and worker
462
- attempt is reconciled to confirmed settlement and each unfinished candidate
463
- is preserved; the old worktree is removed; the old run directory is made
464
- read-only; the new pair is enabled at distinct minutes; the first real driver
465
- tick is recorded. A failure leaves the new pair disabled, and both pairs never
466
- run together. Historical artefacts are untouched.
467
-
468
- ## Exclusions
469
-
470
- No obligations table, supersede counter, or review-budget hold. No `COMMENT`
471
- reviews. No watch deadline or `expired` state. No terminal hygiene or global
472
- busy guard. No terminal nudge from Hermes. No Hermes access to `gh`, `git`, or
473
- `orca`. No re-review of human-placed blocks. No per-project state. No changes
474
- to retired artefacts. No gate for health findings.
1
+ # Native PR managers
2
+
3
+ Read this only for the two native Orca PR-manager automations. The historical
4
+ automation specs and plans describe retired designs and are not instructions.
5
+
6
+ ## Topology and schedules
7
+
8
+ There are exactly two reusable manager chats:
9
+
10
+ - **Review manager:** runs at minutes `0,15,30,45` and invokes
11
+ [axstack-review](../../axstack-review/SKILL.md) for eligible peer reviews.
12
+ - **Watch manager:** runs at minutes `7,22,37,52` and invokes
13
+ [axstack-watch](../../axstack-watch/SKILL.md) for eligible own-PR watch or
14
+ repair events.
15
+
16
+ Each uses native existing-workspace reuse in its own dedicated manager
17
+ workspace. A manager never checks out a PR branch in that workspace. Missed
18
+ slots do not replay a backlog; the next ordinary tick discovers current state.
19
+ The two short packaged prompts sit beside this file and discover these rules by
20
+ relative link instead of copying them.
21
+
22
+ This is prompt policy, not proof that Orca serializes a busy tick or reuses a
23
+ session. Before activation a native canary must prove same-session reuse, busy
24
+ tick behavior, recovery after session loss, nested dispatch depth for
25
+ coordinator-launched leaves, and total process and memory effects. A firing
26
+ timestamp proves neither delivery nor useful completion.
27
+
28
+ ## Discovery and coverage
29
+
30
+ Resolve self on every tick with `gh api user --jq .login`; never hardcode the
31
+ account. Read every discovery page. If pagination or an API call fails, report
32
+ coverage incomplete and make no completeness claim; never silently cap the
33
+ monitored set.
34
+
35
+ The review manager covers open non-draft PRs across accessible repositories
36
+ that are authored by someone else and either officially request review from
37
+ self or have a non-self comment that explicitly mentions self and requests a
38
+ review or response. Incidental mentions grant no authority. Preserve a prior
39
+ human `CHANGES_REQUESTED` block across head changes. It has workflow provenance
40
+ only when its body ends with
41
+ `<!-- axstack-automation verdict head=<sha> -->` bound to its reviewed commit,
42
+ or a retained legacy receipt proves that provenance; otherwise never replace
43
+ it automatically.
44
+
45
+ The watch manager covers every open non-draft PR authored by self. It may
46
+ repair only `defi-com/monorepo` and `defi-com/mobile`; PRs elsewhere remain
47
+ observed but read-only. Preserve deploy-on-push exclusions, lowest-first stack
48
+ dependencies, and existing ownership boundaries. Peer code is always
49
+ read-only.
50
+
51
+ Coverage is not execution. Waiting for CI, a reviewer, a user decision, or a
52
+ merge occupies no execution slot after owned work and descendants settle.
53
+ Watch membership never reserves a slot and no job stays active merely until a
54
+ PR merges or closes. Thirty open PRs, including ten settled waiting PRs, are
55
+ all scanned; those ten occupy zero slots.
56
+
57
+ ## Admission and fairness
58
+
59
+ Each manager admits at most five concurrently executing PR tasks across ticks.
60
+ The review manager admits at most five, and the watch manager admits at most
61
+ five. Their caps are separate; never borrow unused capacity from the other
62
+ lane. Admit fewer when the whole worker tree would put the host under resource
63
+ pressure.
64
+
65
+ A slot covers one bounded PR event and remains occupied while its author,
66
+ reviewers, or other owned descendants are active or unsettled. Leaf workers do
67
+ not create recursive teams. Settlement of the PR job and every descendant
68
+ frees the slot even while the PR stays open.
69
+
70
+ Inspect all eligible PRs before admission. Preserve unserved work in the
71
+ compact run record and select the oldest actionable unserved event first, with
72
+ ascending repository and PR-number tie breaks. A sixth event is admitted after
73
+ a slot settles; repeatedly changing PRs cannot starve older unserved work.
74
+
75
+ ## Per-PR jobs
76
+
77
+ The reusable manager owns ongoing discovery and continuity; the bounded PR
78
+ coordinator owns only its admitted event. Do not create a second live owner or
79
+ writer for the same PR. Reuse an existing valid per-PR worktree, owner, and
80
+ unchanged receipts before creating anything. Otherwise create one separate
81
+ Orca worktree per PR job, parented to that repository's primary worktree, and
82
+ pin the observed head and base. The bounded PR coordinator loads the
83
+ appropriate skill, launches only the reviewers or author that skill owns,
84
+ handles the current actionable event, returns exact receipts, then settles.
85
+ Settlement returns continuity to the manager rather than retaining an idle PR
86
+ coordinator. Reviewers retain the isolation required by `axstack-review`.
87
+
88
+ An unchanged exact head and event creates no job. Dedupe from current GitHub
89
+ state, native Orca Task and Dispatch state, and the existing compact run record;
90
+ do not create machine cursor files or a queue engine. Record enough to resume:
91
+ PR, head, base, event identity, mode, owner and worker receipts, candidate,
92
+ publication receipt, hold, and next action. GitHub remains authoritative for
93
+ open state, revisions, reviews, checks, and merge state.
94
+
95
+ When the current event is handled, settle and release owned native resources.
96
+ Preserve a dirty worktree, an unpushed candidate, pending external result, or
97
+ user-owned work until its durability and ownership are proven. Here a pending
98
+ external result means an unconfirmed publication or send outcome, not pending
99
+ CI. Waiting state belongs in GitHub and the compact record, never in an idle
100
+ model, per-PR timer, or polling loop.
101
+
102
+ ## Review and repair authority
103
+
104
+ Peer review follows `axstack-review` peer mode: two isolated configured
105
+ reviewers inspect the exact head and base. Complete review may publish the
106
+ ordinary binding `APPROVE` or `REQUEST_CHANGES` verdict after a final head,
107
+ base, request, open-state, and existing-review readback. Incomplete review,
108
+ unknown GitHub state, unavailable required models, or unresolved disagreement
109
+ publishes nothing.
110
+
111
+ The public verdict is bound to the GitHub review commit parameter, ends with
112
+ the workflow marker above, and is read back by review ID at that head. It is
113
+ self-contained for the PR reader: include every validated finding and its
114
+ evidence and consequence; never narrate reviewer counts, gates, receipts, or
115
+ private or local artifacts. Write the local HTML copy under the established
116
+ `~/defi/misc/reviews/review-<repo>-PR-<num>.html` convention, with
117
+ `review-PR-<num>.html` reserved for `defi-com/monorepo`. Never publish a
118
+ `COMMENT` review. An ambiguous submission is looked up before retry.
119
+
120
+ An own PR becomes actionable for repair only when a failing check has a base
121
+ check-run with the same name and producing app identity observed passing (or a
122
+ legacy status has the same context), or when a new current-head
123
+ `CHANGES_REQUESTED` review has both an unhandled review ID and an unhandled body
124
+ digest. Missing, pending, or same-name/different-app base evidence holds repair.
125
+ A new head or generic event grants no repair authority by itself.
126
+
127
+ Authorized own-PR repair follows `axstack-watch`: produce the smallest repair
128
+ in the per-PR worktree, obtain the actual-author-provenance reviewer pairing,
129
+ and publish only a reviewed fast-forward repair push after exact-current
130
+ candidate, head, base, event, allowlist, deploy, and remote readback receipts.
131
+ Reconcile an ambiguous review or push result before any retry.
132
+
133
+ No manager, coordinator, or worker may merge, close, force-push, rebase,
134
+ restack, broaden scope, or use `gh stack` mutation. Human merge remains the
135
+ boundary.
136
+
137
+ ## Exceptional decisions and notifications
138
+
139
+ A credible security concern, permanent on-chain state change, or architecture
140
+ decision is held in the reusable manager's Orca conversation. Store the PR,
141
+ head, base, action, candidate, decision context, and preserved candidate bytes
142
+ or refs in the compact record. Send one deduplicated Telegram notification only
143
+ when the recorded `Notification policy` authorizes it, using
144
+ [axstack-relay](../../axstack-relay/SKILL.md) and telling the user to act in
145
+ that manager conversation.
146
+
147
+ Telegram delivery, a Telegram reply, or silence never authorizes an action.
148
+ After a decision in the manager conversation, revalidate the exact candidate,
149
+ head, base, event, authority, and remote state before acting. A changed input
150
+ makes the old decision stale and holds that action. There are no token files,
151
+ Telegram decision interpreter, or separate model gate.
152
+
153
+ ## Recovery and limits
154
+
155
+ On a lost manager session, native recovery first reconciles actual Orca
156
+ workers and Dispatches, GitHub state, and the compact run record. Reuse valid
157
+ unchanged receipts. Unknown ownership blocks only the affected PR, as does
158
+ unknown liveness, approval, or publication outcome; recovery never copies old
159
+ capability, replaces a live writer, or takes over live user work. Other
160
+ unambiguous work may proceed.
161
+
162
+ Use only native schedules and Orca orchestration. Add no daemon, shell precheck,
163
+ watchdog script, custom scheduler, cursor or pending sidecar, runtime database,
164
+ workflow state machine, decision interpreter, or programmatic escalation gate.
165
+ The live VPS activation, native reuse and busy-tick behavior, recovery path,
166
+ nested dispatch depth for coordinator-launched leaves, and resource ceiling
167
+ remain unverified until the canary succeeds.