axstack 0.9.1 → 0.11.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -107,12 +107,11 @@ answer trust or permission prompts on the worker's behalf, and never create a
107
107
  duplicate writer. A `worker_done` advances work only when its Task and Dispatch
108
108
  match the active attempt and its revision evidence verifies.
109
109
 
110
- The user lifted the native-watch hold by decision on 2026-09-16. Orca's native automation
111
- schema exposes a provider but cannot pin model, effort, or permission, so the
112
- driver automation records its model identity every tick and the watchdog treats
113
- a mismatch as a safety hold. Axstack uses no historical fallback and introduces
114
- no custom scheduler. The five-minute driver, hourly read-only watchdog, quiet
115
- healthy ticks, and shared 24-hour deadline remain the acceptance contract.
110
+ The user lifted the native-watch hold by decision on 2026-09-16. The driver
111
+ every 15 minutes dispatches and exits; the watchdog is model-free and
112
+ read-only, has no gate, and records `watchdog.log`; there is no watch deadline
113
+ for automations. Axstack uses no historical fallback and introduces no custom
114
+ scheduler.
116
115
 
117
116
  Mobile completion/reply behavior remains unverified. Structural checks and
118
117
  qualitative scenario evaluation are not live runtime proof.
package/docs/workflows.md CHANGED
@@ -153,44 +153,43 @@ keeps the current owner and a resumable record.
153
153
  Serious security, downtime, data-loss, and major-design risks are raised in a
154
154
  prompt immediately and hold dependent dangerous work. This is not a runtime
155
155
  gate. An applicable `Notification policy` may use `axstack-relay`; otherwise the
156
- current Orca conversation is the fallback. The relay delivers one-way through
157
- native `hermes send`: it checks CLI lookup and the configured target, binds the
158
- recipient, deduplicates on the run record, records the returned `message_id`,
159
- and treats Telegram replies as neither receipts nor authority. Delivery failure
160
- never clears the underlying hold.
156
+ current Orca conversation is the fallback. The relay normally delivers one-way
157
+ through native `hermes send`: it checks CLI lookup and the configured target,
158
+ binds the recipient, deduplicates on the run record, and records the returned
159
+ `message_id`. The PR automation's decision tokens are the narrow exception: a
160
+ fixed Hermes script writes the user's bound decision to a file for the driver
161
+ to consume; no session polls Telegram. Delivery failure never clears the
162
+ underlying hold.
161
163
 
162
164
  Healthy watch observations remain quiet. The optional `axstack-monitor` is
163
- read-only and never sends; `axstack-watchdog` never mutates GitHub and performs
164
- exactly one kind of send, a gate-authorized automation-health escalation
165
- recorded in `watchdog.json`.
165
+ read-only and never sends; `axstack-watchdog` never mutates GitHub. Under the
166
+ rev-3 PR automation it performs four model-free health checks and sends new
167
+ occurrences directly, with no health gate.
166
168
 
167
169
  ## Native watch automations
168
170
 
169
- The accepted monitoring contract is a five-minute driver automation, an hourly
170
- watchdog, quiet healthy snapshots, deduplicated actionable events, verified
171
- handshakes, one owner, and one shared default 24-hour deadline.
171
+ The accepted PR-automation contract is a 15-minute driver automation and an
172
+ hourly watchdog, with quiet healthy checks, deduplicated occurrences, one live
173
+ dispatch marker per PR, and decision tokens for user-authorized actions.
172
174
 
173
- The user lifted the native-watch hold by user decision on 2026-09-16. The driver
174
- automation is a mutating owner for the PRs it handles; the watchdog keeps the
175
- independent read-only contract. The driver is the automation session itself,
176
- with no `axstack-monitor` or `axstack-owner` role row. Native Orca automations still select only a
177
- provider, so the driver records its model identity every tick and the watchdog
178
- treats a mismatch as a safety hold. Axstack adds no custom scheduler, polling
179
- loop, or historical runtime fallback.
175
+ The driver is the automation session itself, with no `axstack-monitor` or
176
+ `axstack-owner` role row. It validates its Opus identity every tick and holds
177
+ instead of dispatching on mismatch. Orca owns scheduling and run history;
178
+ Axstack adds no custom scheduler, polling loop, or historical runtime fallback.
180
179
 
181
180
  ## Automations
182
181
 
183
- Two native Orca automations run the installed skills without a human in the
184
- loop: a driver every five minutes that discovers own and peer PRs, repairs own
185
- PRs in the mutation allowlist, and reviews peer PRs; and a read-only
186
- watchdog every hour that reports automation health. After every
187
- mode-required reviewer settles, the `axstack-auditor` gate returns
188
- exactly one token, `escalate` or `proceed`.
189
- `escalate` records and notifies a hold and publishes nothing.
190
- Only `proceed` plus no unresolved validated blocking finding permits publication
191
- (a fast-forward push or one `COMMENT` review). The approved contract is
192
- `docs/specs/orca-automations.md`; the skill-facing restatement an automation
193
- session loads is `skills/axstack/references/automations.md`.
182
+ Two native Orca automations run the installed skills: a driver every 15 minutes
183
+ that discovers work through four GitHub searches, dispatches exact-head peer
184
+ reviews and own-PR repairs, consumes decision tokens, and exits; and an hourly
185
+ model-free watchdog that evaluates four liveness checks and exits without
186
+ launching a session. PR reviewers use exactly three escalation criteria. A gate
187
+ `escalate` opens a bound decision token and exits, while `proceed` permits only
188
+ the verdict or fast-forward push supported by the reviewed evidence. There are
189
+ no `COMMENT` reviews, obligations, watch deadline, terminal cleanup sweep, or
190
+ health gate. The approved contract is `docs/specs/pr-automations.md` revision 3;
191
+ the installed restatement is
192
+ `skills/axstack/references/automations.md`.
194
193
 
195
194
  ## Run record and evidence
196
195
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "axstack",
3
- "version": "0.9.1",
3
+ "version": "0.11.1",
4
4
  "description": "Axstack installer and setup CLI: installs owned chat skills and role data, configures supported harness settings, and checks Orca capabilities.",
5
5
  "keywords": [
6
6
  "claude-code",
@@ -1,368 +1,342 @@
1
1
  # Automation sessions
2
2
 
3
- Read this when the current session is an Orca automation running a PR driver or
4
- a watchdog.
5
-
6
- ## Pair identity
7
-
8
- There are two independent automation pairs. Before acting, resolve which pair
9
- this session belongs to from its own run id and the verbatim allowlist in its
10
- own prompt; never assume.
11
-
12
- - **Pair A/B** — driver Automation A (`*/30`) and watchdog Automation B, run id
13
- `20260916-pr-automations`, allowlist `axatbhardwaj/axstack`.
14
- - **Pair C/D** — driver Automation C (hourly) and watchdog Automation D, run id
15
- `20260916-defi-automations`, allowlist
16
- `defi-com/monorepo, defi-com/mobile, defi-com/azure-next-hybrid`.
17
-
18
- The two pairs share no state: each has its own run id, run directory,
19
- `progress.md` and sidecars, and no sidecar is shared between them. A clause
20
- below that names a pair applies only to that pair; every other clause applies to
21
- both. Where this reference says "the driver" it means the driver of the current
22
- pair. It restates the operational contract of the approved
23
- `docs/specs/orca-automations.md`; that spec is authoritative, and nothing here
24
- widens it. Orca owns scheduling, sessions, retries, and run history. Axstack
25
- owns policy, the run record, and evidence.
26
-
27
- ## Identity and scope
28
-
29
- - **Self** is `gh api user --jq .login`, resolved at the start of every run and
30
- never hardcoded.
31
- - **Own PR:** an open PR authored by self. Authority equals the user at the
32
- keyboard: repair, test, commit, and fast-forward push to the PR branch
33
- (`git push`, no lease or force). `gh stack` sync or restack is out of scope.
34
- - **Peer PR:** an open PR where self is officially review-requested, or where a
35
- PR comment @-mentions self asking for a response. Unsolicited reviews never
36
- happen. Peer code is read-only. A mention counts as a peer request only when
37
- the driver reads the comment and it explicitly asks self to review or
38
- respond. An incidental mention is discovery data and never review authority.
39
- - **Discovery** covers every repository the account can see:
40
- `gh search prs --state open --limit 100` with `--author @me`,
41
- `--review-requested @me`, and `--mentions @me`. A result count equal to the
42
- limit is a detectable truncation and is logged as `error`. Results are
43
- deduplicated canonically by PR URL across the three searches with own-PR
44
- precedence.
45
- - **Mutation allowlist:** stated verbatim in the automation prompt, per
46
- automation rather than as one global list. One allowlist gates both own-PR
47
- repair and peer review for its own automation; there is no separate
48
- review-only list. Outside the
49
- allowlist the automation discovers and records only. The precheck writes the
50
- full discovery list to `pending.json`; no review, watch, gate, or other model
51
- work is launched for an external PR. External changes do not enter the wake
52
- fingerprint, so external churn never wakes the session. External PR contents
53
- never produce an escalation. This is a deliberate coverage reduction, not an
54
- implied clean review; discovery errors themselves remain automation-health
55
- findings.
56
- - **Prohibited everywhere:** force-push, rebase, merge, close, `APPROVE`,
57
- `REQUEST_CHANGES`. A need for any of these becomes a recorded hold.
58
-
59
- ## Roles per mode
60
-
61
- The automation session is the driver and the owner for every PR it handles;
62
- the driver is the automation session itself, with no `axstack-monitor` or
63
- `axstack-owner` role row materialized, and the standalone-owner branches
64
- of `axstack-watch` and `axstack-review` do not fire. `axstack-monitor` stays
65
- an optional read-only observer that never sends.
66
-
67
- - Peer PR → peer mode of [axstack-review](../../axstack-review/SKILL.md):
68
- `axstack-reviewer-primary` and `axstack-reviewer-secondary`, isolated.
69
- - Own-PR repair → authored mode: the repair author is the automation session
70
- (provider `claude`), so the one cross-family reviewer comes from actual
71
- provenance (`axstack-reviewer-primary`, Sol, in `mixed`). A repair delegated
72
- to `axstack-author` (Sol) takes `axstack-reviewer-secondary`.
73
- - "Every mode-required reviewer" means both reviewers in peer mode and the one
74
- selected reviewer in authored mode.
75
- - Gate → `axstack-auditor`.
76
- - Watchdog → `axstack-watchdog`, which never mutates GitHub and performs
77
- exactly one kind of send, a gate-authorized automation-health escalation
78
- recorded in `watchdog.json`.
79
-
80
- ## Reviewer brief and criteria
81
-
82
- Every reviewer brief ends with a required field, exactly:
3
+ Read this when the current session is the native Orca PR **driver** or its
4
+ **watchdog**. The approved contract is
5
+ `docs/specs/pr-automations.md` revision 4; this reference restates the parts an
6
+ automation session must execute and does not widen them.
7
+
8
+ Pair A/B is retired for this contract; its artefacts remain untouched.
9
+
10
+ ## Roles and authority
11
+
12
+ - **driver** — every 15 minutes, fresh Opus session in the dedicated Axstack
13
+ automations worktree. It discovers GitHub work, binds the persistent Orca
14
+ Run, creates one project-local child worktree per selected PR, dispatches the
15
+ matching Axstack agent, reconciles completions and decisions, then exits. The
16
+ driver is the automation session itself, with no `axstack-monitor` or
17
+ `axstack-owner` role row. It never performs review or repair in its own
18
+ session. Its only direct GitHub mutation is an already-approved decision:
19
+ one `gh pr review` or one fast-forward push of a preserved candidate.
20
+ - **watchdog** — hourly at a distinct minute. It is a shell precheck with no
21
+ model, never launches an agent session, never mutates GitHub, and only sends
22
+ the health notifications defined below.
23
+ - `axstack-monitor` remains an optional read-only observer that never sends.
24
+
25
+ Self is resolved on every run with `gh api user --jq .login`; never hardcode
26
+ it. The review allowlist is `defi-com/monorepo`, `defi-com/mobile`, and
27
+ `defi-com/azure-next-hybrid`. The repair allowlist is `defi-com/monorepo` and
28
+ `defi-com/mobile`. State both lists verbatim in the driver prompt. An own PR is
29
+ open, authored by self, on the repair allowlist, and not a draft. A peer PR is
30
+ open, on the review allowlist, authored by someone else, and either officially
31
+ review-requested from self or has a non-self comment that both mentions self
32
+ and asks for review or response. The driver reads and records the comment id
33
+ and its interpretation; incidental mentions are discovery only.
34
+
35
+ Outside the union of both allowlists, record only and exclude the PR from the
36
+ fingerprint. Peer code is read-only, though its worktree may install
37
+ dependencies and run repository tests. Own-PR authority permits repair, test,
38
+ commit, and fast-forward `git push` without lease or force. Force-push, rebase,
39
+ merge, close, every `gh stack` sync/restack/rebase/merge/link/submit action,
40
+ `COMMENT` reviews, and every GitHub write by the watchdog or Hermes are
41
+ prohibited.
42
+
43
+ ## Discovery and wake
44
+
45
+ The bounded driver precheck contains no model and exits 0 only for changed or
46
+ due work.
47
+
48
+ 1. If `cursor.json.tick_started_at` is newer than `tick_done_at` and younger
49
+ than 1 h, append `running` to `precheck.log` and exit 3. This overlap guard
50
+ inspects no terminal.
51
+ 2. Resolve self, then run exactly four `gh search prs --state open --limit 100
52
+ --json url,number,repository,updatedAt` searches: `--author @me`,
53
+ `--review-requested @me`, `--mentions @me`, and
54
+ `--reviewed-by @me --review changes_requested`. A result count equal to 100
55
+ is truncation: append `error`, exit 2, and do not write a fingerprint.
56
+ Deduplicate by URL with own-PR precedence and drop repositories outside the
57
+ allowlist union.
58
+ 3. For each retained PR, read `gh pr view --json headRefOid,baseRefName,isDraft,
59
+ statusCheckRollup,author,latestReviews` and resolve the base SHA with
60
+ `gh api repos/<repo>/commits/<base>`. Own PRs contribute head, base, draft,
61
+ checks reduced to `{name, conclusion|state}` pairs, and the set of
62
+ `latestReviews` `{id, state}` whose `commit` is the head; others contribute
63
+ head and base.
64
+ 4. Debounce peer and fourth-search heads: they enter the hashed subset only on
65
+ the second consecutive precheck that observes them. The first observation
66
+ goes in `pending.json.seen[]`. Own heads enter immediately. Hash this subset
67
+ as the fingerprint.
68
+
69
+ ### Due control work
70
+
71
+ Read `cursor.json` and `decisions/`. Work is due for:
72
+
73
+ - a decision in `approved` or `rejected` without `consumed_at`;
74
+ - a `spent` decision without a receipt, which needs reconciliation;
75
+ - a `deferred[]` entry whose head still matches discovery;
76
+ - an expired repair cap;
77
+ - a dispatch marker older than 3 h;
78
+ - an unsettled Orca delivery in `pending_settlement[]`.
79
+
80
+ An `open` decision is not due. Write `pending.json` with the fingerprint,
81
+ `observed_at`, `seen[]`, full discovery list, and hashed subset. Append
82
+ `<ts> changed|due|unchanged|running|error` to `precheck.log`; exit 0 for
83
+ `changed` or `due`, 1 for `unchanged`, 2 for `error`, and 3 for `running`.
84
+
85
+ ## Driver tick
86
+
87
+ The driver performs this order and exits:
88
+
89
+ 1. If `tick_started_at` is newer than `tick_done_at`, add `previous tick did
90
+ not finish` to `cursor.json.health[]`. Write `tick_started_at`. Match this
91
+ session's `ORCA_TERMINAL_HANDLE` to the driver's `terminalPtyId` from
92
+ `orca automations runs --id <driver>`, then verify the transcript model.
93
+ Non-Opus or unknown identity records a hold and `tick_outcome: held`, and
94
+ dispatches nothing.
95
+ 2. Bind the persistent Run with `orca orchestration run-use` and read the
96
+ inbox. For a `worker_done` matching a live marker, verify the review id at
97
+ the bound head, push range, or opened token. Release the worker; remove its
98
+ child worktree only after confirmed process exit and accepted settlement;
99
+ then clear the marker. Unverifiable delivery stays in
100
+ `pending_settlement[]` and blocks only that PR.
101
+ 3. Consume decisions as their sole consumer under "Decision tokens" below.
102
+ 4. Reconcile every marker older than 3 h. A live worker gets `worker-stop`; an
103
+ exited worker gets `worker-abandon`. Unknown liveness or user takeover
104
+ retains worktree and marker and blocks only that PR. After confirmed
105
+ abandon, remove the worktree, append one health line, increment the head's
106
+ `abandon_count`, and drop that head so normal selection retries once. For a
107
+ review-triggered dispatch, use the marker's `trigger` to remove exactly its
108
+ review id and digest from `processed_reviews[]`. A second abandon at that
109
+ head is a user-owned hold.
110
+ 5. Select changed PRs and matching `deferred[]` entries oldest `updatedAt`
111
+ first. Skip a current-head decision in `open` or `approved`, and skip a live
112
+ marker. Dispatch within the repair and review rules below.
113
+ 6. Promote the `pending.json` fingerprint verbatim, because it records
114
+ observation rather than completion. Write `tick_done_at` and
115
+ `tick_outcome: ok`, then exit.
116
+
117
+ ## Dispatch and repair selection
118
+
119
+ Every selected PR receives one dispatch marker with task id, dispatch id,
120
+ worktree, head, `started_at`, reservation (`verdict` or `repair`), and trigger:
121
+ `{kind: check, name, app_id}` or `{kind: review, review_id, digest}`. There is at
122
+ most one live marker per PR. The marker is the claim shared by scheduled and
123
+ attended sessions; revalidation immediately before an external call is its
124
+ second half. It makes no exactly-once claim against concurrent human GitHub
125
+ activity.
126
+
127
+ An own PR needs repair when either trigger applies:
128
+
129
+ 1. a failing check has a base check-run with the same `name` and the same
130
+ producing `app.id` observed passing through
131
+ `gh api repos/<repo>/commits/<base>/check-runs`; for a legacy commit status,
132
+ its counterpart has the same `context`. A missing, pending, or same-name
133
+ different-app base check holds repair;
134
+ 2. a `CHANGES_REQUESTED` review at the current head, by any account, has both
135
+ a review id absent from `cursor.json.processed_reviews[]` and a SHA-256 body
136
+ digest not recorded for that PR and head. Both keys are required: the same
137
+ finding under a new review id must not re-trigger repair. Record review id
138
+ and body digest when dispatching. A superseded head with a new review
139
+ triggers again subject to the 24 h cap.
140
+
141
+ Repair also requires no deploy-on-push head branch, no live repair cap, and
142
+ selection of the lowest own PR in its stack that needs repair. Create a child
143
+ worktree in that project's clone at the exact head, parented to the driver
144
+ worktree, and dispatch one `axstack-watch` agent in authored repair mode. Its
145
+ brief contains only the triggering checks or review findings. Each open
146
+ descendant records one user-owned `pending restack` hold until it stops needing
147
+ repair. The 24 h cap starts at dispatch and an abandon does not refund it.
148
+
149
+ A debounced peer PR is eligible when self has not reviewed its head. Read
150
+ `gh pr view --json reviews` before dispatch. Whenever any self review with
151
+ state `CHANGES_REQUESTED` exists on the PR, from whichever search it came,
152
+ dispatch only if the latest effective, non-dismissed self review body contains
153
+ the line prefix `<!-- axstack-automation verdict` or its id is listed in
154
+ `cursor.json.legacy_automation_reviews[]`; otherwise record and skip, so a
155
+ human-placed block is never overwritten. A dismissed block and a self-approved
156
+ PR are skipped. Dispatch one `axstack-review` agent in peer mode and link the
157
+ prior review in its brief.
158
+
159
+ Budgets are one `verdict` dispatch per tick, oldest first; at most six `repair`
160
+ markers live across all repositories; and one repair per PR per 24 h. Put every
161
+ eligible PR not dispatched because of a budget in `deferred[]` with repo, PR,
162
+ and head. Budget exhaustion records the count and is not a hold.
163
+
164
+ ## Agents and verdicts
165
+
166
+ Every agent works in its own project-local Orca child worktree at the exact
167
+ head and reports only through the Orca worker protocol.
168
+
169
+ Peer review runs the two isolated configured reviewers on the identical brief,
170
+ then the Luna gate. `APPROVE` requires complete exact-head/base reviews, gate
171
+ `proceed`, and zero validated blockers. `REQUEST_CHANGES` requires the same
172
+ completeness and `proceed`, plus at least one evidenced blocking finding.
173
+ `INCOMPLETE`, unresolved disagreement, unavailable review or gate, or unknown
174
+ GitHub state publishes nothing. Immediately before `gh pr review`, the agent
175
+ re-reads self's reviews at the head and skips with the existing id when one is
176
+ already present, then re-checks head, base, draft, authorship, and allowlist.
177
+ The verdict body ends with this exact marker line:
83
178
 
84
179
  ```text
85
- Escalate to user: yes | no — <criterion> — <reason>
180
+ <!-- axstack-automation verdict head=<sha> -->
86
181
  ```
87
182
 
88
- Criteria, exactly four: a security concern; a permanent on-chain state change;
89
- an architectural change in approach; and automation health (a
90
- model-substitution, session, precheck, discovery, or relay-delivery hold). The
91
- automation health criterion is usable only by the watchdog and the safety-hold
92
- path, never by a reviewer.
93
-
94
- ## Stack-aware repair sequencing (pair C/D)
95
-
96
- These three sections govern pair C/D only. Pair A/B's behaviour is unchanged by
97
- revision 4 except for the named cadence correction, so A does not apply stack
98
- sequencing, bot-triggered repair, the caps, or the deployment push hold.
99
-
100
- Own-PR repair covers every own PR in the allowlisted repositories; stacking
101
- changes the order of repair, never the scope. Within one stack the driver
102
- repairs only the lowest open failing PR of that stack. Every open descendant of
103
- a repaired PR records exactly one `pending restack` hold naming the repaired
104
- parent and its new head SHA, and receives no independent repair of the same
105
- finding while that hold stands. A descendant carrying a different finding from
106
- the repaired parent's is repaired on its own merits, subject to the budgets
107
- below; the hold suppresses duplicate application of the same finding, not all
108
- work on the descendant.
109
-
110
- The hold names the user as owner. It is cleared by the user's own restack, or by
111
- a later tick observing the descendant no longer failing; that is the single
112
- clearing rule. It is exempt from the watchdog's hold with no owner threshold,
113
- because its owner is the user by construction. The per-tick budget counts one
114
- unit per stack, not one per PR.
115
-
116
- The reason is duplication, not politeness: the same finding recurs across a
117
- stack, so independent per-PR repair would apply the identical fix at several
118
- levels and the user's next cascade rebase would then conflict on the duplicate.
119
- A parent fast-forward also does not change a descendant's PR diff, so a
120
- descendant's CI and reviewers never see the parent fix; the hold states that
121
- honestly instead of implying the descendant was repaired.
122
-
123
- `gh stack` sync, restack, rebase, merge, link and submit remain out of scope for
124
- every automation, and force-push and rebase stay prohibited everywhere. A
125
- descendant is held, never rewritten.
126
-
127
- ## Bot review feedback (pair C/D)
128
-
129
- A review authored by a bot account may be actionable and may trigger a repair,
130
- bounded as follows.
131
-
132
- Dedup is by processed review ID and by a digest of the review body. A repair
133
- fires only on a review whose review ID was never processed and whose body
134
- digest is not already recorded against that PR at that head SHA. Review-ID dedup alone is
135
- insufficient: a bot that re-posts the identical finding under a new review ID
136
- would otherwise re-trigger a repair on every tick.
137
-
138
- Caps: at most one repair per PR per 24 hours, counted regardless of trigger; and
139
- a per-tick push budget across all allowlisted repositories combined, configured
140
- per pair and six for pair C/D. Reaching either cap records a hold naming the cap
141
- and the value reached, rather than silently dropping the work. A budget hold
142
- names the budget as its owner and is exempt from the hold with no owner
143
- threshold. When a per-PR cap has expired the precheck wakes the driver even
144
- if the forge is unchanged, so capped work is never stranded; the precheck
145
- section below carries that trigger in its due-work list.
146
-
147
- A bot review never satisfies the peer-PR trigger, which still requires self to
148
- be officially review-requested or a comment that explicitly asks self to review
149
- or respond.
150
-
151
- ## Watch window
152
-
153
- Pair A/B keeps its 24 hours per own PR from first observation, ending early on
154
- merge or close, with the deadline in `cursor.json`, the `expired` marking, and
155
- the user re-arming an expired PR in the automation session.
156
-
157
- Pair C/D uses a rolling window instead: a PR is in scope while it is open and
158
- eligible, and leaves scope on merge or close. There is no expiry and no
159
- re-arming, C/D's `cursor.json` stores no deadlines, and `expired` is not a state
160
- a C/D PR can reach.
161
-
162
- The difference is deliberate, not drift. A serves one low-volume repository
163
- where expiry is cheap, while C would otherwise start twenty or more simultaneous
164
- clocks on first observation and go dark a day later. Because that window
165
- removes expiry as a cost brake, the per-PR and per-tick budgets and the watchdog
166
- are the only brakes left on C, and D's thresholds are retuned for a fingerprint
167
- that legitimately changes on nearly every tick.
168
-
169
- ## PR eligibility for repair (pair C/D)
170
-
171
- A draft PR is discovered and recorded but never repaired; it becomes eligible
172
- when it is marked ready for review, which the ordinary event state observes as
173
- new work. A PR that already has a human reviewer requested is eligible for
174
- repair and is not excluded, deliberately: most own PRs in these repositories
175
- carry a requested human reviewer, and excluding them would empty the coverage.
176
- `defi-com/mobile` has no workflows and therefore no check signal, so a repair
177
- there is triggered only by actionable review feedback; if CI is later added the
178
- ordinary check-rollup trigger applies with no contract change.
179
-
180
- ## Deployment safety (pair C/D)
181
-
182
- A repair never pushes to a PR whose head branch is in that repository's
183
- deploy-on-push set, and an attempt to do so is a recorded hold. The rule is
184
- branch-name-agnostic: the set is enumerated per repository from that
185
- repository's workflow files, recorded, and re-verified before enabling and on
186
- any later allowlist change. It is never hardcoded to one branch name, because a
187
- repository may deploy from more than one branch and may add another at any time.
188
-
189
- ## Escalation gate
190
-
191
- After every mode-required reviewer settles, the driver spawns `axstack-auditor`
192
- with the verdicts and the candidate revision, instructing it to act as the
193
- escalation gate. It returns exactly one literal token, `escalate` or `proceed`.
194
-
195
- - The gate decides only whether the user is notified. Reviewer "yes" is input,
196
- not a veto.
197
- - `escalate` → records the hold, then one `hermes send` through
198
- [axstack-relay](../../axstack-relay/SKILL.md) naming the PR, the criterion,
199
- every reviewer's reason, where the user acts (Orca conversation, worktree, or
200
- PR), and the hold; it publishes nothing.
201
- - `proceed` → no notification; a push or `COMMENT` publication then
202
- additionally requires no unresolved validated blocking finding, because
203
- `proceed` never overrides a validated blocking finding. A reviewer security
204
- "yes" that the gate does not escalate is recorded as rejected-with-evidence
205
- or returned to the author before any mutation.
206
- - Only `proceed` plus no unresolved validated blocking finding permits a push
207
- or publication; a push before the gate settles is forbidden.
208
- - Precedence: credible serious risk found by a reviewer still produces the
209
- standing internal prompt and dependent-action hold immediately, and the gate
210
- governs only external notification. That hold is the serious-risk rule of
211
- [contracts](contracts.md#serious-risk). The internal prompt lands in the run
212
- record and the automation's Orca conversation; no `hermes send` occurs
213
- without `escalate`.
214
- - The health gate takes a watchdog finding and its evidence, not reviewer
215
- verdicts or a candidate; the same two tokens apply.
216
- - An unavailable gate or required reviewer records a hold, pauses mutation for
217
- that PR, and is treated as a watchdog health finding. An unavailable gate
218
- cannot be escalated through itself: the hold stays, the gap is visible in
219
- Orca run history and the sidecar, and no substitute or unauthorized send
220
- occurs.
221
-
222
- ## Precheck (bounded shell, no model)
223
-
224
- Resolve self; run the three searches with `--json url,number,repository,updatedAt`;
225
- write the full discovery list to `pending.json`. For each allowlisted,
226
- non-expired own PR add head SHA, base SHA, and the check rollup of that head via
227
- `gh pr view --json headRefOid,baseRefOid,statusCheckRollup`; check state is
228
- data, and only authentication, command, and network errors are `error`. For pair
229
- C/D only, that call also requests `isDraft` and the precheck adds draft status
230
- to the hashed fingerprint, so a draft becoming ready wakes the driver on its
231
- own; pair A/B's queried fields and fingerprint are unchanged. Hash
232
- only allowlisted PRs. Read `cursor.json` for the last processed fingerprint and
233
- for due control work: a watch deadline at or before now (pair A/B only, since
234
- pair C/D stores no deadlines), a pending failed-relay retry, or a per-PR repair
235
- cap that has expired (pair C/D only, since pair A/B has no caps). Exit 0 when the hash differs or control work is due;
236
- otherwise exit non-zero. Exit non-zero without running when the previous driver
237
- run is still active, or on `error`. Append one line
238
- `<ts> <changed|due|unchanged|error|busy>` to `precheck.log`.
239
-
240
- Terminal hygiene runs before the searches and covers driver terminals only.
241
- Ownership is the driver automation's own recorded `terminalPtyId` from its Orca
242
- run history, never a terminal title: Orca rewrites a Claude terminal's title to
243
- the agent's current task summary, so a driver's title drifts and an unrelated
244
- session can acquire one that reads like a driver. The precheck closes the
245
- previous ticks' idle driver terminals with `--tab` — without it the pane closes
246
- but the session stays listed and is never reclaimed, which also makes an
247
- over-match destructive — and a driver terminal that is still working makes the
248
- tick `busy`. It never closes a watchdog terminal or any terminal outside this
249
- automation; an unreadable ownership source is `error`, never a silent empty
250
- sweep. No two of the four automations may share a dispatch minute: A, B, C and D each
251
- take a distinct minute, so a driver and its watchdog never collide and neither
252
- pair can disturb the other's terminal hygiene.
253
-
254
- The observed fingerprint is written to `pending.json`. After the processed
255
- tick the driver promotes exactly that value to `cursor.json`, never a
256
- recomputed one, so an event landing during a run is processed on the following
257
- tick.
183
+ Write the local review file to the workspace review directory
184
+ `~/defi/misc/reviews/` under the existing convention:
185
+ `review-PR-<num>.html` with no prefix means `defi-com/monorepo`;
186
+ `review-mobile-PR-<num>.html` and
187
+ `review-azure-next-hybrid-PR-<num>.html` name the other repositories. A write
188
+ failure is recorded but does not withhold the verdict.
258
189
 
259
- ## Driver tick
190
+ Authored repair commits a local candidate, obtains one Sol review at that
191
+ local SHA and the Luna gate, resolves every validated blocker, records test
192
+ evidence, re-reads remote head/base/draft/deploy set/allowlist, then pushes
193
+ fast-forward. The monorepo unit gate uses `~/.bun-1.2.2/bin/bun`; a suite that
194
+ cannot run locally is an explicit unverified boundary.
260
195
 
261
- Before any repair or publication the driver validates its effective session
262
- identity through Orca runtime inspection and records it; a self-written label
263
- is not evidence. Expected model: Opus. A mismatch or unknown identity holds
264
- repair and publication for that run and is a health finding.
265
-
266
- For each changed PR:
267
-
268
- - Own PR → [axstack-watch](../../axstack-watch/SKILL.md) on the exact head
269
- SHA in a per-PR child worktree; the driver worktree never checks out a PR
270
- branch. Follow its repair-publication reference: candidate committed locally,
271
- authored review at the local SHA, gate, then fast-forward push with the
272
- publication readback immediately before it.
273
- - Peer PR → two isolated `axstack-review` passes on the exact head SHA, then
274
- the gate, then one owner-synthesized `COMMENT` review under the review skill's
275
- COMMENT branch. `INCOMPLETE` or an unavailable required reviewer records a
276
- hold and publishes nothing.
277
-
278
- Watch window, pair A/B only: 24 hours per own PR from first observation, ending
279
- early on merge or close. The deadline is stored in `cursor.json`; a due deadline
280
- wakes the driver through the precheck even when GitHub is unchanged, and the
281
- driver rechecks the deadline immediately before any publication. Expiry marks
282
- the PR `expired` in the record and sidecar, records a resumable handoff, and
283
- stops silently. The driver skips an expired PR until the user re-arms it in the
284
- automation session, and the precheck ignores it; an expired PR is never silently
285
- re-adopted. Pair C/D does not use this window at all; see "Watch window" above
286
- for its rolling replacement, and it stores no deadline and reaches no `expired`
287
- state.
288
-
289
- Per-PR event state: head SHA, base SHA, check rollup, and processed request and
290
- comment IDs. A review receipt is reused only when head, base, and scope are
291
- unchanged. A new failing check, base change, review request, or qualifying
292
- comment at an unchanged head is new work. Pair C/D additionally carries draft
293
- status in that state, and a draft status that has changed from true to false is
294
- new work for C/D, which is how a draft becoming ready for review reaches its
295
- driver.
296
-
297
- ## Watchdog tick
298
-
299
- Before its gate dispatch the watchdog validates its own effective session
300
- identity through Orca runtime inspection; unknown identity holds the dispatch
301
- and is itself recorded in `watchdog.json`.
302
-
303
- Read Orca run history for the driver, `precheck.log`, and the run record.
304
- Thresholds: three consecutive `error` lines in `precheck.log`; three
305
- consecutive failed driver runs; no successful driver run within two hours while
306
- the precheck logged `changed` or `due`; any unrequested fresh-session fallback
307
- or non-Opus effective identity recorded by the driver; any `failed` relay
308
- receipt older than one tick or any `uncertain` receipt; a hold with no owner.
309
- A quiet precheck history with no due work is healthy.
310
-
311
- The two-hour stall threshold above is pair A/B's. Pair D uses a stall window
312
- re-derived before enabling as `max(2h, 3 x the 95th-percentile observed C tick
313
- duration over at least 10 ticks)`, recorded with its sample, because C's
314
- fingerprint legitimately changes on nearly every tick and a literal two hours
315
- would fire on the first slow tick. Every other threshold is identical for both
316
- pairs.
317
-
318
- Each health finding gets an occurrence id `(type, first-observed UTC
319
- timestamp)`; it stays deduplicated while unresolved, and a later recurrence is
320
- a new occurrence. Pass a finding to the gate under the automation-health
321
- criterion. On `escalate` the watchdog itself performs that one gate-authorized
322
- `hermes send` and records the receipt in `watchdog.json`; it never mutates
323
- GitHub and never writes `progress.md` or `cursor.json`. Failed or uncertain
324
- delivery stays visibly held in `watchdog.json` and Orca run history.
325
-
326
- ## Run record and sidecar
327
-
328
- One run id per pair for the lifetime of that pair — `20260916-pr-automations`
329
- for A/B and `20260916-defi-automations` for C/D — each in the
330
- [run record](run-record.md) shape with that pair's driver as sole writer of its
331
- own `progress.md`. C/D's run directory lives under the same axstack
332
- `git-common-dir` as A/B's, in its own `axstack/runs/<run id>/` folder; no
333
- sidecar, record, or cursor is shared between the pairs. Per PR it stores processed event IDs, exact head and base SHAs, review receipts
334
- per SHA, gate decisions, `hermes send` receipts with `message_id` and state, and
335
- holds. The watch deadline and `expired` fields are pair A/B only, since pair C/D
336
- stores no deadline and cannot reach `expired`; draft status is pair C/D only,
337
- since only C/D treats a draft transition as new work. Its
338
- `Notification policy:` line reads, verbatim:
196
+ Every reviewer brief ends exactly:
339
197
 
340
198
  ```text
341
- hermes send, target telegram (home), host VPS, gate-authorized escalations only
199
+ Escalate to user: yes | no — <criterion> — <reason>
342
200
  ```
343
201
 
344
- Machine-readable sidecars in the same directory: `pending.json` (observed
345
- fingerprint plus full discovery list, written by the precheck), `cursor.json`
346
- (last processed fingerprint promoted verbatim from `pending.json`, expired PR
347
- list and per-PR watch deadlines for pair A/B only, pending failed-relay
348
- retries; written by the driver after each processed tick), `precheck.log` (precheck only), and
349
- `watchdog.json` (watchdog only). Orca run history remains the authoritative
350
- log; the record is derived progress, never authority.
351
-
352
- Dedup: a processed event ID is never processed twice; a review receipt is
353
- reused only for an unchanged head, base, and scope. PR notifications dedup on
354
- (PR, criterion, head SHA); health notifications dedup on the occurrence id. A
355
- `failed` relay receipt may be retried once, on the next tick, as due control
356
- work; an `uncertain` one is never auto-resent.
202
+ PR work has exactly three criteria: a security concern, a permanent on-chain
203
+ state change, or an architectural change in approach. There is no
204
+ automation-health criterion for reviewers. The Luna gate returns exactly one
205
+ token, `escalate` or `proceed`; there is no gate for health findings.
206
+ `escalate` opens a decision token, sends its message, and exits without waiting
207
+ for a reply. `proceed` does not override a validated blocker. `worker_done`
208
+ names the PR, head, action, GitHub receipt or opened token. The driver alone
209
+ writes the run record.
210
+
211
+ ## Decision tokens
212
+
213
+ Store `decisions/<token>.json`, where `<token>` contains at least 96 random
214
+ bits as lowercase hex, produced for example by `openssl rand -hex 16`.
215
+ Immutable bound fields are written once: `repo`, `pr`, `head`, `base`,
216
+ `action` (`approve-verdict`, `request-changes-verdict`, or `push`), plus exact
217
+ `body` and `commit` for verdicts or `candidate_sha`, `head_branch`, and
218
+ `expected_remote_head` for a push. Mutable fields are `state`, `created_at`,
219
+ `decided_at`, `decided_message_id`, `consumed_at`, `receipt`, `send`, and
220
+ `reason`. Never delete a token file.
221
+
222
+ Lifecycle has one named writer per transition, each by temp file + rename:
223
+
224
+ | transition | writer |
225
+ | --- | --- |
226
+ | create `open` | the agent that escalated |
227
+ | `open → approved` / `open → rejected` | the Hermes script only |
228
+ | `approved → spent` / `approved → stale` | the driver only |
229
+ | `rejected → closed` | the driver only |
230
+ | `spent` + `receipt` | the driver only |
231
+
232
+ ### Opening
233
+
234
+ Before opening a `push` token, the repair agent pins its candidate with local
235
+ ref `refs/axstack/decisions/<token>` in the project clone, so the candidate
236
+ survives worktree removal. It then sends one
237
+ `hermes send --to telegram` message naming the PR, criterion, every reviewer's
238
+ reason, and the exact replies `approve <token>` and `reject <token>`. Store the
239
+ send receipt. The driver retries a `failed` send once next tick and reconciles
240
+ an `uncertain` send against Hermes output before any retry.
241
+
242
+ The Hermes gateway's fixed `axstack-decide` script accepts only those two exact
243
+ commands. It validates the private user/channel configuration on every call,
244
+ restricts tokens to `^[0-9a-f]{24,}$`, locks and re-reads an `open` token,
245
+ writes only its decision fields by temp file and rename, and prints one line.
246
+ Hermes never runs `gh`, `git`, or `orca`.
247
+
248
+ ### Consuming
249
+
250
+ The driver is the only consumer. Immediately before acting it re-reads that
251
+ `state == approved`, then revalidates open/unmerged state, bound head/base,
252
+ absence of a self review for a verdict, or expected remote head plus reachable
253
+ candidate for a push. It writes `spent` before the GitHub call, executes
254
+ exactly the bound action with no second gate, then stores the receipt. A spent
255
+ token without a receipt is
256
+ reconciliation: match the exact review commit/body or destination
257
+ ref/candidate on GitHub; record a match, or the driver retries once under the
258
+ same approval after proving non-execution, or hold ambiguity. A revalidation
259
+ failure writes `stale`
260
+ with the reason and drops the PR's head from `cursor.json`; the driver never
261
+ mints a token. Close rejected files and skip that head. After `spent` or
262
+ `stale`, delete the candidate ref. Tokens do not expire; the watchdog reports
263
+ one open longer than 24 h.
264
+
265
+ ## Watchdog
266
+
267
+ The hourly shell precheck only reads `cursor.json`, `precheck.log`,
268
+ `decisions/`, and `orca automations runs --id <driver>`. It performs exactly
269
+ these four checks and always exits non-zero, so no model session launches:
270
+
271
+ | check | trips when |
272
+ | --- | --- |
273
+ | driver stuck | a `changed` or `due` precheck is older than 1 h with no later `tick_done_at`, including a driver that never wrote `tick_started_at` (Orca run status alone is not evidence of completion) |
274
+ | precheck failing | the last three `precheck.log` entries are `error` |
275
+ | worker stuck | a dispatch marker is older than 3 h 30 min and remains uncleared |
276
+ | decision waiting | an `open` decision is older than 24 h |
277
+
278
+ Each trip is `(check, first_observed)`. Write one JSON line per tick to
279
+ `watchdog.log` with `{ts, checks, trips, sent}` and send each occurrence once
280
+ through `hermes send --to telegram`. Persist `sent`, `failed`, or `uncertain`;
281
+ retry `failed` next tick and reconcile `uncertain` before retry. Re-send only
282
+ after the check was observed clear and later recurs. Unreadable evidence is an
283
+ `unknown` occurrence. A healthy tick writes its line and sends nothing. There
284
+ is no gate for health findings.
285
+
286
+ ## Run directory
287
+
288
+ One `.git/axstack/runs/<run id>/` directory under the Axstack git-common-dir
289
+ contains:
290
+
291
+ - `cursor.json` — driver only, with these exact keys: `fingerprint`,
292
+ `tick_started_at`, `tick_done_at`, `tick_outcome`, `prs{url: {head, base,
293
+ draft, checks, reviews, last_self_review}}`, `dispatch_markers[]` (`pr`,
294
+ `task_id`, `dispatch_id`, `worktree`, `head`, `started_at`, `reservation`,
295
+ `trigger`), `deferred[]`, `pending_settlement[]`,
296
+ `repair_caps{url: {expires_at}}`, `abandon_count{head: n}`,
297
+ `processed_reviews[]` (`review_id`, `pr`, `head`, `digest`),
298
+ `deploy_on_push{repo: [branches]}`,
299
+ `legacy_automation_reviews[]`, `health[]`;
300
+ - `pending.json`, `precheck.log` — driver precheck only;
301
+ - `decisions/<token>.json` — writers assigned by the lifecycle table;
302
+ - `watchdog.log` and `watchdog-state.json` (occurrence `first_observed` values
303
+ and send receipts) — watchdog only;
304
+ - `progress.md` — driver only, one line per tick plus holds and mention
305
+ readings, with no per-PR prose.
306
+
307
+ Timestamps are UTC `YYYY-MM-DDTHH:MM:SSZ`; an unparsable timestamp is an
308
+ `error` for the precheck and `unknown` for the watchdog, never silently
309
+ ignored.
310
+
311
+ Orca run history is the authoritative log. The launch worktree is a dedicated
312
+ Axstack worktree where nobody develops. The current `defi-automations`
313
+ worktree retires with the old pair. Keep this notification-policy edge in the
314
+ run record: `Notification policy` authorizes the token and watchdog sends;
315
+ delivery uses [axstack-relay](../../axstack-relay/SKILL.md).
357
316
 
358
317
  ## Safety holds
359
318
 
360
- - The driver records its effective identity on every tick. An unrequested
361
- fresh-session fallback, or a recorded identity different from the expected
362
- one, pauses mutation for that run, is recorded, and goes to the watchdog
363
- path. A user-run `--fresh-session` reconciles from the run record and is not
364
- a hold.
365
- - GitHub API errors leave the PR state unknown; nothing is pushed or published
366
- on unknown state.
367
- - A hold is cleared only by a later run observing the condition resolved, or by
368
- the user in the Orca conversation. Silence never clears a hold.
319
+ - Non-Opus or unknown driver identity records a hold and dispatches nothing.
320
+ - A GitHub API error makes PR state unknown; never publish or push for it.
321
+ - Enumerate deploy-on-push branches from both repair repositories before
322
+ enabling and store them in `cursor.json`; re-check on allowlist changes.
323
+ - A later tick observing resolution or an explicit user decision clears a
324
+ hold. Silence never clears one.
325
+
326
+ ## Cutover
327
+
328
+ Perform this order: the new pair exists disabled; the amended skills and
329
+ references are installed; C and D are disabled; every old driver and worker
330
+ attempt is reconciled to confirmed settlement and each unfinished candidate
331
+ is preserved; the old worktree is removed; the old run directory is made
332
+ read-only; the new pair is enabled at distinct minutes; the first real driver
333
+ tick is recorded. A failure leaves the new pair disabled, and both pairs never
334
+ run together. Historical artefacts are untouched.
335
+
336
+ ## Exclusions
337
+
338
+ No obligations table, supersede counter, or review-budget hold. No `COMMENT`
339
+ reviews. No watch deadline or `expired` state. No terminal hygiene or global
340
+ busy guard. No terminal nudge from Hermes. No Hermes access to `gh`, `git`, or
341
+ `orca`. No re-review of human-placed blocks. No per-project state. No changes
342
+ to retired artefacts. No gate for health findings.
@@ -17,12 +17,9 @@ binding state and receipts to exact revisions.
17
17
  `axstack-reviewer-secondary` sessions with identical brief and isolated first
18
18
  pass; authored = one eligible configured reviewer from actual author
19
19
  provenance. Owner and author never review.
20
- - Driver/monitor/watchdog: the five-minute driver is the automation session
21
- itself, a mutating owner for the PRs it handles with no `axstack-monitor` or
22
- `axstack-owner` role row. `axstack-monitor` is an optional read-only observer
23
- that never sends. `axstack-watchdog` is the hourly health checker that never
24
- mutates GitHub and performs exactly one kind of send, a gate-authorized
25
- automation-health escalation recorded in `watchdog.json`.
20
+ - Driver/monitor/watchdog: the driver every 15 minutes dispatches and exits as
21
+ a mutating owner; the watchdog is model-free and read-only, has no gate, and
22
+ records `watchdog.log`; there is no watch deadline for automations.
26
23
  - Auditor (`axstack-auditor`): report-only; never edits, merges, activates, or
27
24
  audits itself.
28
25
 
@@ -99,19 +96,18 @@ Tracking grants no merge, release, model-substitution, or scope authority.
99
96
 
100
97
  ## Deadline (one rule for every owned timer)
101
98
 
102
- The default 24-hour deadline covers every task-owned timer, including open-PR
103
- monitors and watchdogs. Stop at deadline; remaining work gets a resumable
104
- handoff, never silent renewal. Merge-ready differs from merged; human merges.
99
+ The default 24-hour deadline covers standalone task-owned timers. Stop them at
100
+ deadline and preserve remaining work; there is no watch deadline for
101
+ automations. Merge-ready differs from merged; human merges.
105
102
 
106
103
  ## Watch health
107
104
 
108
- The user lifted the native-watch hold by user decision on 2026-09-16. The driver
109
- automation is a mutating owner for its PRs; `axstack-watchdog` stays
110
- independent and read-only, with quiet healthy snapshots, deduplicated events,
111
- restart reconciliation, and one shared deadline. Native Orca automations expose
112
- a provider but cannot pin model, effort, or permission, so the driver records
113
- its model identity every tick and the watchdog treats a mismatch as a safety
114
- hold. Build no custom scheduler and use no legacy fallback. Details live in
105
+ The user lifted the native-watch hold by user decision on 2026-09-16. The
106
+ driver automation is a mutating owner for its PRs; `axstack-watchdog` stays
107
+ independent and read-only. The driver every 15 minutes dispatches and exits;
108
+ the watchdog is model-free, has no gate, and records `watchdog.log`; there is
109
+ no watch deadline for automations. Build no custom scheduler and use no legacy
110
+ fallback. Details live in
115
111
  [Watch runtime](../../axstack-watch/references/watch-runtime.md).
116
112
 
117
113
  ## Audit hook (end of run and meaningful checkpoints)
@@ -63,12 +63,12 @@ listing all pass.
63
63
 
64
64
  ## Preserve identity and authority
65
65
 
66
- Delivery is one-way. Hermes does not route a Telegram reply back to the
67
- sending session; its own agent answers replies. A reply is therefore never a
68
- receipt, decision, or authority for this session, and no persistent owner is
69
- needed to send. Every message must say where the user acts: the current Orca
70
- conversation, the Orca worktree, or the GitHub PR. Do not ask the user to reply with
71
- decision words, and do not poll Telegram for answers.
66
+ Delivery is one-way; no session polls Telegram. Hermes does not route a reply
67
+ back to the sending session; its own agent answers replies. Outside the fixed
68
+ decision-token flow below, a reply is never a receipt, decision, or authority
69
+ for this session, and no persistent owner is needed to send. Every ordinary
70
+ message must say where the user acts: the current Orca conversation, the Orca
71
+ worktree, or the GitHub PR. Do not invent reply commands.
72
72
 
73
73
  Send authority comes from the explicit request or applicable standing policy.
74
74
  It grants no merge, publication, ownership-transfer, or model-substitution
@@ -76,6 +76,24 @@ authority. Delivery is transport evidence only. Revalidate any user decision
76
76
  that arrives through an authorized channel against the current task and
77
77
  existing action boundaries before acting; silence never grants permission.
78
78
 
79
+ ## Decision tokens
80
+
81
+ An automation escalation is the narrow exception defined by
82
+ [Automation sessions](../axstack/references/automations.md). The PR agent opens
83
+ an immutable-bound decision token, sends one message with the exact
84
+ `approve <token>` and `reject <token>` replies, records the send receipt, and
85
+ exits without waiting. The Hermes script decides by validating the private
86
+ channel and updating only an open token file; Hermes never performs the bound
87
+ GitHub or Git action. The driver consumes the file on a later tick, revalidates
88
+ all bound state, marks an approval spent before acting, performs only that
89
+ action, and records its receipt.
90
+
91
+ This does not create a reply channel for the sending agent: Hermes pushes the
92
+ decision to the file through the fixed `axstack-decide` script, and the driver
93
+ reads that file. Delivery is one-way; no session polls Telegram. Token creation,
94
+ decision, and consumption keep their separate writers and authority; a message
95
+ receipt alone authorizes nothing.
96
+
79
97
  ## Reconcile, deliver, and record
80
98
 
81
99
  Before any new send, check the caller's run record for an existing receipt
@@ -16,10 +16,10 @@ to select the mode and scope identity, and apply the shared
16
16
  [PR-shape policy](../axstack/references/pr-shape.md). For an owned candidate,
17
17
  load and verify the
18
18
  [candidate-publication boundary](../axstack/references/candidate-publication.md).
19
- When the caller is the Orca driver automation, load
20
- [Automation sessions](../axstack/references/automations.md): its reviewer
21
- briefs carry the required escalation field and its publication is `COMMENT`
22
- only.
19
+ When the caller is an Orca driver automation, load
20
+ [Automation sessions](../axstack/references/automations.md): its reviewer briefs
21
+ carry the required escalation field and every eligible peer PR takes a binding
22
+ `APPROVE` or `REQUEST_CHANGES` verdict under the automation exception below.
23
23
 
24
24
  ## Peer mode (colleague PR)
25
25
 
@@ -219,11 +219,10 @@ Escalate to user: yes | no — <criterion> — <reason>
219
219
 
220
220
  Every brief ends with the `Escalate to user` field and the reviewer answers it
221
221
  in the receipt. A reviewer may cite only a security concern, a permanent
222
- on-chain state change, or an architectural change in approach; the automation
223
- health criterion belongs to the watchdog and safety-hold path and is never a
224
- reviewer criterion. The answer is input to the escalation gate, not a veto and
225
- not a verdict; see [Automation sessions](../axstack/references/automations.md)
226
- for the gate.
222
+ on-chain state change, or an architectural change in approach. Health is not a
223
+ reviewer criterion. The answer is input to the PR escalation gate, not a veto
224
+ and not a verdict; see
225
+ [Automation sessions](../axstack/references/automations.md) for the gate.
227
226
 
228
227
  ## Template: review receipt (one block per revision)
229
228
 
@@ -258,9 +257,9 @@ private transport values or configuration.
258
257
  Under an automation session, credible serious risk found by a reviewer still
259
258
  raises the standing internal prompt and dependent-action hold immediately, and
260
259
  the gate governs only external notification: the internal prompt lands in the
261
- run record and the automation session's own Orca conversation, and no
262
- `hermes send` occurs without the gate's `escalate` token. `proceed` never
263
- overrides a validated blocking finding.
260
+ run record and the automation session's own Orca conversation. `escalate`
261
+ opens the bound decision token, sends through `axstack-relay`, and exits
262
+ without waiting; `proceed` never overrides a validated blocking finding.
264
263
 
265
264
  ## Publishing rule
266
265
 
@@ -306,38 +305,26 @@ for a complete `APPROVE` or `REQUEST_CHANGES` verdict:
306
305
  Submission is complete only when the remote receipt confirms the review bound
307
306
  to the intended commit.
308
307
 
309
- Automation exception — Authorized submission: the `COMMENT` branch below, with
310
- the existing remote head/base readback and ambiguity handling, is the only
311
- submission the Orca driver automation makes. The prohibition on `APPROVE` and
312
- `REQUEST_CHANGES` is a ban on those GitHub actions; the review skill's internal
313
- verdict vocabulary is unchanged.
314
-
315
- ## Automation publication (`COMMENT`)
316
-
317
- Only the Orca driver automation, as owner for a peer PR under
318
- [Automation sessions](../axstack/references/automations.md), uses this branch:
319
- one `COMMENT` review, owner-synthesized and bound to the reviewed commit. It
320
- never submits `APPROVE` or `REQUEST_CHANGES`; a need for either is a
321
- recorded hold. The peer-review submission rule above is unchanged for every
322
- other caller.
323
-
324
- 1. Complete the mode-required review: every mode-required receipt is current
325
- for the head SHA and current base, and each carries its escalation field.
326
- `INCOMPLETE`, unresolved material disagreement, or an unavailable required
327
- reviewer publishes nothing and records a hold that pauses mutation for the
328
- PR.
329
- 2. The gate returns exactly one token, `escalate` or `proceed`. `escalate`
330
- records the hold and the relay receipt and publishes nothing. Only
331
- `proceed` plus no unresolved validated blocking finding permits
332
- publication.
333
- 3. Read back the remote head and base immediately before submit; stop if
334
- either differs from the reviewed candidate.
335
- 4. Submit one owner-synthesized `COMMENT` review bound to the reviewed commit
336
- through the actual GitHub commit parameter, carrying every receipt's
337
- findings and limitations, then verify the submission receipt.
338
- 5. A receipt for an unchanged head SHA is never published twice; on ambiguity,
339
- look up remote state and submit only when absent.
340
-
341
- Publication is complete when the remote receipt confirms one `COMMENT` review
342
- bound to the intended commit, or when the recorded hold names the missing
343
- input and the next owner.
308
+ ## Automation exception
309
+
310
+ For a peer PR selected under
311
+ [Automation sessions](../axstack/references/automations.md), apply the same
312
+ complete-review and exact-commit requirements, then run the Luna gate.
313
+ `escalate` opens a bound decision token, sends its message, and exits without
314
+ submitting. `proceed` permits `APPROVE` only with no validated blocker and
315
+ permits `REQUEST_CHANGES` only with at least one evidenced validated blocker.
316
+ `INCOMPLETE`, unavailable inputs, unresolved disagreement, or unknown GitHub
317
+ state submits nothing.
318
+
319
+ Immediately before `gh pr review`, re-read self's reviews at the head. If one
320
+ already exists, skip submission and record its id. Otherwise re-check head,
321
+ base, draft status, authorship, and allowlist, bind the verdict to the exact
322
+ commit parameter, and end the body with this marker line:
323
+
324
+ ```text
325
+ <!-- axstack-automation verdict head=<sha> -->
326
+ ```
327
+
328
+ After submission, read back and record the review id at the bound head. An
329
+ ambiguous result is looked up before any retry. The automation never submits a
330
+ `COMMENT` review.
@@ -22,7 +22,7 @@ The authenticated human login selects PRs. Runtime session IDs coordinate work
22
22
  only and establish neither human identity nor write, reply, or merge authority.
23
23
  When the session is the Orca driver automation, also load
24
24
  [Automation sessions](../axstack/references/automations.md): it is the owner
25
- for every PR it handles, and its gate and allowlist bound every mutation.
25
+ for every PR it handles, and its allowlist bounds every mutation.
26
26
 
27
27
  ## 1. Adopt and reconcile
28
28
 
@@ -63,9 +63,10 @@ Read-only checks and updates to the already-owned local record need no runtime
63
63
  load. When the watch needs a new owner or automated observation, first read
64
64
  [Watch runtime](references/watch-runtime.md) and then
65
65
  [Orca runtime](../axstack/references/orca-runtime.md). Reconcile before creating
66
- anything. The user lifted the native-watch hold by user decision: the 5 min driver
67
- automation is a mutating owner for the PRs it handles and the hourly watchdog
68
- stays independent and read-only. The driver is the automation session itself,
66
+ anything. The user lifted the native-watch hold by user decision: the driver
67
+ every 15 minutes dispatches and exits as a mutating owner; the watchdog is
68
+ model-free and read-only, has no gate, and records `watchdog.log`; there is no
69
+ watch deadline for automations. The driver is the automation session itself,
69
70
  with no `axstack-monitor` or `axstack-owner` role row; `axstack-monitor` stays
70
71
  an optional read-only observer that never sends. One read-only PR observation
71
72
  needs neither.
@@ -114,10 +115,10 @@ the current revision, and a recorded hold or next owner where work remains.
114
115
 
115
116
  When a new actionable event is eligible under a recorded `Notification policy`,
116
117
  the owner may use the optional [axstack-relay](../axstack-relay/SKILL.md).
117
- The monitor never sends, and `axstack-watchdog` never mutates GitHub and
118
- performs exactly one kind of send, a gate-authorized automation-health
119
- escalation recorded in `watchdog.json`; absent policy or failed relay uses the
120
- current Orca conversation and leaves every existing hold open.
118
+ The monitor never sends. The automation watchdog never mutates GitHub, has no
119
+ gate, and records occurrences and send receipts in `watchdog.log`; absent
120
+ policy or failed relay uses the current Orca conversation and leaves every
121
+ existing hold open.
121
122
 
122
123
  ## 5. State readiness precisely
123
124
 
@@ -128,15 +129,13 @@ observed state distinct from merged, and the human merges by default.
128
129
 
129
130
  ## 6. End and preserve continuity
130
131
 
131
- End early when all required PRs merge, or at cancel or the shared default 24h
132
- deadline. In every case, stop and verify all owned registrations. The deadline
133
- also stops timers for open PRs; never silently renew them.
132
+ End a standalone watch early when all required PRs merge, at cancellation, or
133
+ at its shared default 24 h deadline. There is no watch deadline for
134
+ automations. In every case, stop and verify all owned registrations.
134
135
 
135
136
  At every end condition, leave the compact state below in the private run record
136
137
  and report it in the current chat, even when work remains. Expiry grants neither
137
- silent renewal nor ownership-transfer authority. Under an automation, expiry
138
- marks the PR `expired` in the record and sidecar; an `expired` PR is never
139
- silently re-adopted and is skipped until the user re-arms it.
138
+ silent renewal nor ownership-transfer authority.
140
139
 
141
140
  Transfer ownership through the runtime-owned Orca handoff route only when the
142
141
  user explicitly requests it. Before transfer, follow the lifecycle-owned
@@ -36,9 +36,11 @@ exact new revision and base, all six angles, applicable acceptance, the reply
36
36
  body identities, and every affected boundary, with no unresolved material
37
37
  finding or urgent hold. Under an automation the escalation gate of
38
38
  [Automation sessions](../../axstack/references/automations.md) also runs on
39
- the local SHA: publication additionally requires the gate to return `proceed`
40
- with no unresolved validated blocking finding, and a push before the gate
41
- settles is forbidden.
39
+ the local SHA. `escalate` creates a token, pins the candidate at
40
+ `refs/axstack/decisions/<token>`, opens the decision token, sends its message,
41
+ and exits instead of pushing. `proceed` permits publication only with no
42
+ unresolved validated blocking finding, and a push before the gate settles is
43
+ forbidden.
42
44
 
43
45
  ## 3. Revalidate immediately before publication
44
46
 
@@ -4,18 +4,16 @@ Read this before starting, resuming, or stopping automated PR observation.
4
4
 
5
5
  ## Accepted policy
6
6
 
7
- The PR owner remains accountable throughout one shared default 24-hour window.
8
- `axstack-monitor` and `axstack-watchdog` are independent, read-only roles, not
9
- authors, reviewers, repliers, or owners.
7
+ The PR owner remains accountable throughout a standalone watch's default
8
+ 24-hour window. `axstack-monitor` and `axstack-watchdog` are independent,
9
+ read-only roles, not authors, reviewers, repliers, or owners.
10
10
 
11
11
  - **Monitor:** `axstack-monitor` is an optional read-only observer that reads
12
- GitHub, all PR feedback, and latest checks every five minutes, persists event
12
+ GitHub, all PR feedback, and latest checks on its registered cadence, persists event
13
13
  IDs, wakes the owner only for a new actionable event, and never sends.
14
- - **Watchdog:** `axstack-watchdog` reads only automation health, handshake
15
- state, and snapshot freshness hourly and never mutates GitHub; it may perform
16
- exactly one kind of send, a gate-authorized automation-health escalation
17
- recorded in `watchdog.json`, and otherwise reports a verified health failure
18
- to the owner.
14
+ - **Watchdog:** for the PR automations, it is a model-free hourly shell precheck
15
+ that never mutates GitHub, has no gate, and records each tick in
16
+ `watchdog.log`.
19
17
 
20
18
  Healthy observations are snapshot-only and update quietly; they wake neither owner nor
21
19
  driver. Both roles deduplicate event IDs. Uncertain delivery is reconciled
@@ -32,25 +30,21 @@ permission pinning are unsupported, and its schedule parser cannot preserve
32
30
  the accepted bounded expiry by itself. Requested role values or a post-launch
33
31
  self-report are not effective launch evidence.
34
32
 
35
- The user lifted the native-watch hold by user decision on 2026-09-16. The accepted
36
- contract now has a new shape: the driver automation is a mutating owner for the
37
- PRs it handles, not an independent read-only monitor, and the watchdog keeps
38
- the independent read-only health contract. The driver is the automation
39
- session itself, with no `axstack-monitor` or `axstack-owner` role row
40
- materialized for it. The driver records its own model
41
- identity on every tick and the watchdog compares it with the expected model; a
42
- mismatch is a safety hold, never a silent substitution. The bounded expiry is
43
- enforced by the run record's watch deadline, not by the schedule parser. Still
33
+ The user lifted the native-watch hold by user decision on 2026-09-16. The
34
+ driver every 15 minutes dispatches and exits as the mutating owner, and the
35
+ watchdog is model-free and read-only, has no gate, and records `watchdog.log`;
36
+ there is no watch deadline for automations. The driver is the automation
37
+ session itself, with no `axstack-monitor` or `axstack-owner` role row. Still
44
38
  introduce no custom scheduler or polling loop and use no legacy runtime
45
39
  fallback. The session-level contract lives in
46
40
  [Automation sessions](../../axstack/references/automations.md).
47
41
 
48
42
  ## Preserve the contract under automation
49
43
 
50
- An automation session preserves the roles, five-minute/hourly cadences, quiet
51
- healthy behavior, deduplication, handshake, watched scope, wake owner, and
52
- shared expiry. Test active expiry, missed final ticks, restart, duplicate
53
- ticks, cancellation, session-reuse fallback, and final cleanup before enabling.
44
+ An automation session preserves its 15-minute driver and hourly model-free
45
+ watchdog cadences, quiet healthy behavior, deduplication, handshake, watched
46
+ scope, and wake owner. Test restart, duplicate ticks, cancellation,
47
+ session-reuse fallback, and final cleanup before enabling.
54
48
  A firing timestamp proves neither delivery nor work advancement. An unrequested
55
49
  fallback session reconciles ownership and never becomes owner silently.
56
50