axstack 0.9.1 → 0.11.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +5 -6
- package/docs/workflows.md +28 -29
- package/package.json +1 -1
- package/skills/axstack/references/automations.md +327 -353
- package/skills/axstack/references/lifecycle.md +12 -16
- package/skills/axstack-relay/SKILL.md +24 -6
- package/skills/axstack-review/SKILL.md +34 -47
- package/skills/axstack-watch/SKILL.md +13 -14
- package/skills/axstack-watch/references/repair-publication.md +5 -3
- package/skills/axstack-watch/references/watch-runtime.md +16 -22
package/README.md
CHANGED
|
@@ -107,12 +107,11 @@ answer trust or permission prompts on the worker's behalf, and never create a
|
|
|
107
107
|
duplicate writer. A `worker_done` advances work only when its Task and Dispatch
|
|
108
108
|
match the active attempt and its revision evidence verifies.
|
|
109
109
|
|
|
110
|
-
The user lifted the native-watch hold by decision on 2026-09-16.
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
healthy ticks, and shared 24-hour deadline remain the acceptance contract.
|
|
110
|
+
The user lifted the native-watch hold by decision on 2026-09-16. The driver
|
|
111
|
+
every 15 minutes dispatches and exits; the watchdog is model-free and
|
|
112
|
+
read-only, has no gate, and records `watchdog.log`; there is no watch deadline
|
|
113
|
+
for automations. Axstack uses no historical fallback and introduces no custom
|
|
114
|
+
scheduler.
|
|
116
115
|
|
|
117
116
|
Mobile completion/reply behavior remains unverified. Structural checks and
|
|
118
117
|
qualitative scenario evaluation are not live runtime proof.
|
package/docs/workflows.md
CHANGED
|
@@ -153,44 +153,43 @@ keeps the current owner and a resumable record.
|
|
|
153
153
|
Serious security, downtime, data-loss, and major-design risks are raised in a
|
|
154
154
|
prompt immediately and hold dependent dangerous work. This is not a runtime
|
|
155
155
|
gate. An applicable `Notification policy` may use `axstack-relay`; otherwise the
|
|
156
|
-
current Orca conversation is the fallback. The relay delivers one-way
|
|
157
|
-
native `hermes send`: it checks CLI lookup and the configured target,
|
|
158
|
-
recipient, deduplicates on the run record, records the returned
|
|
159
|
-
|
|
160
|
-
|
|
156
|
+
current Orca conversation is the fallback. The relay normally delivers one-way
|
|
157
|
+
through native `hermes send`: it checks CLI lookup and the configured target,
|
|
158
|
+
binds the recipient, deduplicates on the run record, and records the returned
|
|
159
|
+
`message_id`. The PR automation's decision tokens are the narrow exception: a
|
|
160
|
+
fixed Hermes script writes the user's bound decision to a file for the driver
|
|
161
|
+
to consume; no session polls Telegram. Delivery failure never clears the
|
|
162
|
+
underlying hold.
|
|
161
163
|
|
|
162
164
|
Healthy watch observations remain quiet. The optional `axstack-monitor` is
|
|
163
|
-
read-only and never sends; `axstack-watchdog` never mutates GitHub
|
|
164
|
-
|
|
165
|
-
|
|
165
|
+
read-only and never sends; `axstack-watchdog` never mutates GitHub. Under the
|
|
166
|
+
rev-3 PR automation it performs four model-free health checks and sends new
|
|
167
|
+
occurrences directly, with no health gate.
|
|
166
168
|
|
|
167
169
|
## Native watch automations
|
|
168
170
|
|
|
169
|
-
The accepted
|
|
170
|
-
watchdog, quiet healthy
|
|
171
|
-
|
|
171
|
+
The accepted PR-automation contract is a 15-minute driver automation and an
|
|
172
|
+
hourly watchdog, with quiet healthy checks, deduplicated occurrences, one live
|
|
173
|
+
dispatch marker per PR, and decision tokens for user-authorized actions.
|
|
172
174
|
|
|
173
|
-
The
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
provider, so the driver records its model identity every tick and the watchdog
|
|
178
|
-
treats a mismatch as a safety hold. Axstack adds no custom scheduler, polling
|
|
179
|
-
loop, or historical runtime fallback.
|
|
175
|
+
The driver is the automation session itself, with no `axstack-monitor` or
|
|
176
|
+
`axstack-owner` role row. It validates its Opus identity every tick and holds
|
|
177
|
+
instead of dispatching on mismatch. Orca owns scheduling and run history;
|
|
178
|
+
Axstack adds no custom scheduler, polling loop, or historical runtime fallback.
|
|
180
179
|
|
|
181
180
|
## Automations
|
|
182
181
|
|
|
183
|
-
Two native Orca automations run the installed skills
|
|
184
|
-
|
|
185
|
-
|
|
186
|
-
watchdog
|
|
187
|
-
|
|
188
|
-
|
|
189
|
-
|
|
190
|
-
|
|
191
|
-
|
|
192
|
-
|
|
193
|
-
|
|
182
|
+
Two native Orca automations run the installed skills: a driver every 15 minutes
|
|
183
|
+
that discovers work through four GitHub searches, dispatches exact-head peer
|
|
184
|
+
reviews and own-PR repairs, consumes decision tokens, and exits; and an hourly
|
|
185
|
+
model-free watchdog that evaluates four liveness checks and exits without
|
|
186
|
+
launching a session. PR reviewers use exactly three escalation criteria. A gate
|
|
187
|
+
`escalate` opens a bound decision token and exits, while `proceed` permits only
|
|
188
|
+
the verdict or fast-forward push supported by the reviewed evidence. There are
|
|
189
|
+
no `COMMENT` reviews, obligations, watch deadline, terminal cleanup sweep, or
|
|
190
|
+
health gate. The approved contract is `docs/specs/pr-automations.md` revision 3;
|
|
191
|
+
the installed restatement is
|
|
192
|
+
`skills/axstack/references/automations.md`.
|
|
194
193
|
|
|
195
194
|
## Run record and evidence
|
|
196
195
|
|
package/package.json
CHANGED
|
@@ -1,368 +1,342 @@
|
|
|
1
1
|
# Automation sessions
|
|
2
2
|
|
|
3
|
-
Read this when the current session is
|
|
4
|
-
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
- **
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
`
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
`
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
-
|
|
74
|
-
|
|
75
|
-
-
|
|
76
|
-
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
3
|
+
Read this when the current session is the native Orca PR **driver** or its
|
|
4
|
+
**watchdog**. The approved contract is
|
|
5
|
+
`docs/specs/pr-automations.md` revision 4; this reference restates the parts an
|
|
6
|
+
automation session must execute and does not widen them.
|
|
7
|
+
|
|
8
|
+
Pair A/B is retired for this contract; its artefacts remain untouched.
|
|
9
|
+
|
|
10
|
+
## Roles and authority
|
|
11
|
+
|
|
12
|
+
- **driver** — every 15 minutes, fresh Opus session in the dedicated Axstack
|
|
13
|
+
automations worktree. It discovers GitHub work, binds the persistent Orca
|
|
14
|
+
Run, creates one project-local child worktree per selected PR, dispatches the
|
|
15
|
+
matching Axstack agent, reconciles completions and decisions, then exits. The
|
|
16
|
+
driver is the automation session itself, with no `axstack-monitor` or
|
|
17
|
+
`axstack-owner` role row. It never performs review or repair in its own
|
|
18
|
+
session. Its only direct GitHub mutation is an already-approved decision:
|
|
19
|
+
one `gh pr review` or one fast-forward push of a preserved candidate.
|
|
20
|
+
- **watchdog** — hourly at a distinct minute. It is a shell precheck with no
|
|
21
|
+
model, never launches an agent session, never mutates GitHub, and only sends
|
|
22
|
+
the health notifications defined below.
|
|
23
|
+
- `axstack-monitor` remains an optional read-only observer that never sends.
|
|
24
|
+
|
|
25
|
+
Self is resolved on every run with `gh api user --jq .login`; never hardcode
|
|
26
|
+
it. The review allowlist is `defi-com/monorepo`, `defi-com/mobile`, and
|
|
27
|
+
`defi-com/azure-next-hybrid`. The repair allowlist is `defi-com/monorepo` and
|
|
28
|
+
`defi-com/mobile`. State both lists verbatim in the driver prompt. An own PR is
|
|
29
|
+
open, authored by self, on the repair allowlist, and not a draft. A peer PR is
|
|
30
|
+
open, on the review allowlist, authored by someone else, and either officially
|
|
31
|
+
review-requested from self or has a non-self comment that both mentions self
|
|
32
|
+
and asks for review or response. The driver reads and records the comment id
|
|
33
|
+
and its interpretation; incidental mentions are discovery only.
|
|
34
|
+
|
|
35
|
+
Outside the union of both allowlists, record only and exclude the PR from the
|
|
36
|
+
fingerprint. Peer code is read-only, though its worktree may install
|
|
37
|
+
dependencies and run repository tests. Own-PR authority permits repair, test,
|
|
38
|
+
commit, and fast-forward `git push` without lease or force. Force-push, rebase,
|
|
39
|
+
merge, close, every `gh stack` sync/restack/rebase/merge/link/submit action,
|
|
40
|
+
`COMMENT` reviews, and every GitHub write by the watchdog or Hermes are
|
|
41
|
+
prohibited.
|
|
42
|
+
|
|
43
|
+
## Discovery and wake
|
|
44
|
+
|
|
45
|
+
The bounded driver precheck contains no model and exits 0 only for changed or
|
|
46
|
+
due work.
|
|
47
|
+
|
|
48
|
+
1. If `cursor.json.tick_started_at` is newer than `tick_done_at` and younger
|
|
49
|
+
than 1 h, append `running` to `precheck.log` and exit 3. This overlap guard
|
|
50
|
+
inspects no terminal.
|
|
51
|
+
2. Resolve self, then run exactly four `gh search prs --state open --limit 100
|
|
52
|
+
--json url,number,repository,updatedAt` searches: `--author @me`,
|
|
53
|
+
`--review-requested @me`, `--mentions @me`, and
|
|
54
|
+
`--reviewed-by @me --review changes_requested`. A result count equal to 100
|
|
55
|
+
is truncation: append `error`, exit 2, and do not write a fingerprint.
|
|
56
|
+
Deduplicate by URL with own-PR precedence and drop repositories outside the
|
|
57
|
+
allowlist union.
|
|
58
|
+
3. For each retained PR, read `gh pr view --json headRefOid,baseRefName,isDraft,
|
|
59
|
+
statusCheckRollup,author,latestReviews` and resolve the base SHA with
|
|
60
|
+
`gh api repos/<repo>/commits/<base>`. Own PRs contribute head, base, draft,
|
|
61
|
+
checks reduced to `{name, conclusion|state}` pairs, and the set of
|
|
62
|
+
`latestReviews` `{id, state}` whose `commit` is the head; others contribute
|
|
63
|
+
head and base.
|
|
64
|
+
4. Debounce peer and fourth-search heads: they enter the hashed subset only on
|
|
65
|
+
the second consecutive precheck that observes them. The first observation
|
|
66
|
+
goes in `pending.json.seen[]`. Own heads enter immediately. Hash this subset
|
|
67
|
+
as the fingerprint.
|
|
68
|
+
|
|
69
|
+
### Due control work
|
|
70
|
+
|
|
71
|
+
Read `cursor.json` and `decisions/`. Work is due for:
|
|
72
|
+
|
|
73
|
+
- a decision in `approved` or `rejected` without `consumed_at`;
|
|
74
|
+
- a `spent` decision without a receipt, which needs reconciliation;
|
|
75
|
+
- a `deferred[]` entry whose head still matches discovery;
|
|
76
|
+
- an expired repair cap;
|
|
77
|
+
- a dispatch marker older than 3 h;
|
|
78
|
+
- an unsettled Orca delivery in `pending_settlement[]`.
|
|
79
|
+
|
|
80
|
+
An `open` decision is not due. Write `pending.json` with the fingerprint,
|
|
81
|
+
`observed_at`, `seen[]`, full discovery list, and hashed subset. Append
|
|
82
|
+
`<ts> changed|due|unchanged|running|error` to `precheck.log`; exit 0 for
|
|
83
|
+
`changed` or `due`, 1 for `unchanged`, 2 for `error`, and 3 for `running`.
|
|
84
|
+
|
|
85
|
+
## Driver tick
|
|
86
|
+
|
|
87
|
+
The driver performs this order and exits:
|
|
88
|
+
|
|
89
|
+
1. If `tick_started_at` is newer than `tick_done_at`, add `previous tick did
|
|
90
|
+
not finish` to `cursor.json.health[]`. Write `tick_started_at`. Match this
|
|
91
|
+
session's `ORCA_TERMINAL_HANDLE` to the driver's `terminalPtyId` from
|
|
92
|
+
`orca automations runs --id <driver>`, then verify the transcript model.
|
|
93
|
+
Non-Opus or unknown identity records a hold and `tick_outcome: held`, and
|
|
94
|
+
dispatches nothing.
|
|
95
|
+
2. Bind the persistent Run with `orca orchestration run-use` and read the
|
|
96
|
+
inbox. For a `worker_done` matching a live marker, verify the review id at
|
|
97
|
+
the bound head, push range, or opened token. Release the worker; remove its
|
|
98
|
+
child worktree only after confirmed process exit and accepted settlement;
|
|
99
|
+
then clear the marker. Unverifiable delivery stays in
|
|
100
|
+
`pending_settlement[]` and blocks only that PR.
|
|
101
|
+
3. Consume decisions as their sole consumer under "Decision tokens" below.
|
|
102
|
+
4. Reconcile every marker older than 3 h. A live worker gets `worker-stop`; an
|
|
103
|
+
exited worker gets `worker-abandon`. Unknown liveness or user takeover
|
|
104
|
+
retains worktree and marker and blocks only that PR. After confirmed
|
|
105
|
+
abandon, remove the worktree, append one health line, increment the head's
|
|
106
|
+
`abandon_count`, and drop that head so normal selection retries once. For a
|
|
107
|
+
review-triggered dispatch, use the marker's `trigger` to remove exactly its
|
|
108
|
+
review id and digest from `processed_reviews[]`. A second abandon at that
|
|
109
|
+
head is a user-owned hold.
|
|
110
|
+
5. Select changed PRs and matching `deferred[]` entries oldest `updatedAt`
|
|
111
|
+
first. Skip a current-head decision in `open` or `approved`, and skip a live
|
|
112
|
+
marker. Dispatch within the repair and review rules below.
|
|
113
|
+
6. Promote the `pending.json` fingerprint verbatim, because it records
|
|
114
|
+
observation rather than completion. Write `tick_done_at` and
|
|
115
|
+
`tick_outcome: ok`, then exit.
|
|
116
|
+
|
|
117
|
+
## Dispatch and repair selection
|
|
118
|
+
|
|
119
|
+
Every selected PR receives one dispatch marker with task id, dispatch id,
|
|
120
|
+
worktree, head, `started_at`, reservation (`verdict` or `repair`), and trigger:
|
|
121
|
+
`{kind: check, name, app_id}` or `{kind: review, review_id, digest}`. There is at
|
|
122
|
+
most one live marker per PR. The marker is the claim shared by scheduled and
|
|
123
|
+
attended sessions; revalidation immediately before an external call is its
|
|
124
|
+
second half. It makes no exactly-once claim against concurrent human GitHub
|
|
125
|
+
activity.
|
|
126
|
+
|
|
127
|
+
An own PR needs repair when either trigger applies:
|
|
128
|
+
|
|
129
|
+
1. a failing check has a base check-run with the same `name` and the same
|
|
130
|
+
producing `app.id` observed passing through
|
|
131
|
+
`gh api repos/<repo>/commits/<base>/check-runs`; for a legacy commit status,
|
|
132
|
+
its counterpart has the same `context`. A missing, pending, or same-name
|
|
133
|
+
different-app base check holds repair;
|
|
134
|
+
2. a `CHANGES_REQUESTED` review at the current head, by any account, has both
|
|
135
|
+
a review id absent from `cursor.json.processed_reviews[]` and a SHA-256 body
|
|
136
|
+
digest not recorded for that PR and head. Both keys are required: the same
|
|
137
|
+
finding under a new review id must not re-trigger repair. Record review id
|
|
138
|
+
and body digest when dispatching. A superseded head with a new review
|
|
139
|
+
triggers again subject to the 24 h cap.
|
|
140
|
+
|
|
141
|
+
Repair also requires no deploy-on-push head branch, no live repair cap, and
|
|
142
|
+
selection of the lowest own PR in its stack that needs repair. Create a child
|
|
143
|
+
worktree in that project's clone at the exact head, parented to the driver
|
|
144
|
+
worktree, and dispatch one `axstack-watch` agent in authored repair mode. Its
|
|
145
|
+
brief contains only the triggering checks or review findings. Each open
|
|
146
|
+
descendant records one user-owned `pending restack` hold until it stops needing
|
|
147
|
+
repair. The 24 h cap starts at dispatch and an abandon does not refund it.
|
|
148
|
+
|
|
149
|
+
A debounced peer PR is eligible when self has not reviewed its head. Read
|
|
150
|
+
`gh pr view --json reviews` before dispatch. Whenever any self review with
|
|
151
|
+
state `CHANGES_REQUESTED` exists on the PR, from whichever search it came,
|
|
152
|
+
dispatch only if the latest effective, non-dismissed self review body contains
|
|
153
|
+
the line prefix `<!-- axstack-automation verdict` or its id is listed in
|
|
154
|
+
`cursor.json.legacy_automation_reviews[]`; otherwise record and skip, so a
|
|
155
|
+
human-placed block is never overwritten. A dismissed block and a self-approved
|
|
156
|
+
PR are skipped. Dispatch one `axstack-review` agent in peer mode and link the
|
|
157
|
+
prior review in its brief.
|
|
158
|
+
|
|
159
|
+
Budgets are one `verdict` dispatch per tick, oldest first; at most six `repair`
|
|
160
|
+
markers live across all repositories; and one repair per PR per 24 h. Put every
|
|
161
|
+
eligible PR not dispatched because of a budget in `deferred[]` with repo, PR,
|
|
162
|
+
and head. Budget exhaustion records the count and is not a hold.
|
|
163
|
+
|
|
164
|
+
## Agents and verdicts
|
|
165
|
+
|
|
166
|
+
Every agent works in its own project-local Orca child worktree at the exact
|
|
167
|
+
head and reports only through the Orca worker protocol.
|
|
168
|
+
|
|
169
|
+
Peer review runs the two isolated configured reviewers on the identical brief,
|
|
170
|
+
then the Luna gate. `APPROVE` requires complete exact-head/base reviews, gate
|
|
171
|
+
`proceed`, and zero validated blockers. `REQUEST_CHANGES` requires the same
|
|
172
|
+
completeness and `proceed`, plus at least one evidenced blocking finding.
|
|
173
|
+
`INCOMPLETE`, unresolved disagreement, unavailable review or gate, or unknown
|
|
174
|
+
GitHub state publishes nothing. Immediately before `gh pr review`, the agent
|
|
175
|
+
re-reads self's reviews at the head and skips with the existing id when one is
|
|
176
|
+
already present, then re-checks head, base, draft, authorship, and allowlist.
|
|
177
|
+
The verdict body ends with this exact marker line:
|
|
83
178
|
|
|
84
179
|
```text
|
|
85
|
-
|
|
180
|
+
<!-- axstack-automation verdict head=<sha> -->
|
|
86
181
|
```
|
|
87
182
|
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
## Stack-aware repair sequencing (pair C/D)
|
|
95
|
-
|
|
96
|
-
These three sections govern pair C/D only. Pair A/B's behaviour is unchanged by
|
|
97
|
-
revision 4 except for the named cadence correction, so A does not apply stack
|
|
98
|
-
sequencing, bot-triggered repair, the caps, or the deployment push hold.
|
|
99
|
-
|
|
100
|
-
Own-PR repair covers every own PR in the allowlisted repositories; stacking
|
|
101
|
-
changes the order of repair, never the scope. Within one stack the driver
|
|
102
|
-
repairs only the lowest open failing PR of that stack. Every open descendant of
|
|
103
|
-
a repaired PR records exactly one `pending restack` hold naming the repaired
|
|
104
|
-
parent and its new head SHA, and receives no independent repair of the same
|
|
105
|
-
finding while that hold stands. A descendant carrying a different finding from
|
|
106
|
-
the repaired parent's is repaired on its own merits, subject to the budgets
|
|
107
|
-
below; the hold suppresses duplicate application of the same finding, not all
|
|
108
|
-
work on the descendant.
|
|
109
|
-
|
|
110
|
-
The hold names the user as owner. It is cleared by the user's own restack, or by
|
|
111
|
-
a later tick observing the descendant no longer failing; that is the single
|
|
112
|
-
clearing rule. It is exempt from the watchdog's hold with no owner threshold,
|
|
113
|
-
because its owner is the user by construction. The per-tick budget counts one
|
|
114
|
-
unit per stack, not one per PR.
|
|
115
|
-
|
|
116
|
-
The reason is duplication, not politeness: the same finding recurs across a
|
|
117
|
-
stack, so independent per-PR repair would apply the identical fix at several
|
|
118
|
-
levels and the user's next cascade rebase would then conflict on the duplicate.
|
|
119
|
-
A parent fast-forward also does not change a descendant's PR diff, so a
|
|
120
|
-
descendant's CI and reviewers never see the parent fix; the hold states that
|
|
121
|
-
honestly instead of implying the descendant was repaired.
|
|
122
|
-
|
|
123
|
-
`gh stack` sync, restack, rebase, merge, link and submit remain out of scope for
|
|
124
|
-
every automation, and force-push and rebase stay prohibited everywhere. A
|
|
125
|
-
descendant is held, never rewritten.
|
|
126
|
-
|
|
127
|
-
## Bot review feedback (pair C/D)
|
|
128
|
-
|
|
129
|
-
A review authored by a bot account may be actionable and may trigger a repair,
|
|
130
|
-
bounded as follows.
|
|
131
|
-
|
|
132
|
-
Dedup is by processed review ID and by a digest of the review body. A repair
|
|
133
|
-
fires only on a review whose review ID was never processed and whose body
|
|
134
|
-
digest is not already recorded against that PR at that head SHA. Review-ID dedup alone is
|
|
135
|
-
insufficient: a bot that re-posts the identical finding under a new review ID
|
|
136
|
-
would otherwise re-trigger a repair on every tick.
|
|
137
|
-
|
|
138
|
-
Caps: at most one repair per PR per 24 hours, counted regardless of trigger; and
|
|
139
|
-
a per-tick push budget across all allowlisted repositories combined, configured
|
|
140
|
-
per pair and six for pair C/D. Reaching either cap records a hold naming the cap
|
|
141
|
-
and the value reached, rather than silently dropping the work. A budget hold
|
|
142
|
-
names the budget as its owner and is exempt from the hold with no owner
|
|
143
|
-
threshold. When a per-PR cap has expired the precheck wakes the driver even
|
|
144
|
-
if the forge is unchanged, so capped work is never stranded; the precheck
|
|
145
|
-
section below carries that trigger in its due-work list.
|
|
146
|
-
|
|
147
|
-
A bot review never satisfies the peer-PR trigger, which still requires self to
|
|
148
|
-
be officially review-requested or a comment that explicitly asks self to review
|
|
149
|
-
or respond.
|
|
150
|
-
|
|
151
|
-
## Watch window
|
|
152
|
-
|
|
153
|
-
Pair A/B keeps its 24 hours per own PR from first observation, ending early on
|
|
154
|
-
merge or close, with the deadline in `cursor.json`, the `expired` marking, and
|
|
155
|
-
the user re-arming an expired PR in the automation session.
|
|
156
|
-
|
|
157
|
-
Pair C/D uses a rolling window instead: a PR is in scope while it is open and
|
|
158
|
-
eligible, and leaves scope on merge or close. There is no expiry and no
|
|
159
|
-
re-arming, C/D's `cursor.json` stores no deadlines, and `expired` is not a state
|
|
160
|
-
a C/D PR can reach.
|
|
161
|
-
|
|
162
|
-
The difference is deliberate, not drift. A serves one low-volume repository
|
|
163
|
-
where expiry is cheap, while C would otherwise start twenty or more simultaneous
|
|
164
|
-
clocks on first observation and go dark a day later. Because that window
|
|
165
|
-
removes expiry as a cost brake, the per-PR and per-tick budgets and the watchdog
|
|
166
|
-
are the only brakes left on C, and D's thresholds are retuned for a fingerprint
|
|
167
|
-
that legitimately changes on nearly every tick.
|
|
168
|
-
|
|
169
|
-
## PR eligibility for repair (pair C/D)
|
|
170
|
-
|
|
171
|
-
A draft PR is discovered and recorded but never repaired; it becomes eligible
|
|
172
|
-
when it is marked ready for review, which the ordinary event state observes as
|
|
173
|
-
new work. A PR that already has a human reviewer requested is eligible for
|
|
174
|
-
repair and is not excluded, deliberately: most own PRs in these repositories
|
|
175
|
-
carry a requested human reviewer, and excluding them would empty the coverage.
|
|
176
|
-
`defi-com/mobile` has no workflows and therefore no check signal, so a repair
|
|
177
|
-
there is triggered only by actionable review feedback; if CI is later added the
|
|
178
|
-
ordinary check-rollup trigger applies with no contract change.
|
|
179
|
-
|
|
180
|
-
## Deployment safety (pair C/D)
|
|
181
|
-
|
|
182
|
-
A repair never pushes to a PR whose head branch is in that repository's
|
|
183
|
-
deploy-on-push set, and an attempt to do so is a recorded hold. The rule is
|
|
184
|
-
branch-name-agnostic: the set is enumerated per repository from that
|
|
185
|
-
repository's workflow files, recorded, and re-verified before enabling and on
|
|
186
|
-
any later allowlist change. It is never hardcoded to one branch name, because a
|
|
187
|
-
repository may deploy from more than one branch and may add another at any time.
|
|
188
|
-
|
|
189
|
-
## Escalation gate
|
|
190
|
-
|
|
191
|
-
After every mode-required reviewer settles, the driver spawns `axstack-auditor`
|
|
192
|
-
with the verdicts and the candidate revision, instructing it to act as the
|
|
193
|
-
escalation gate. It returns exactly one literal token, `escalate` or `proceed`.
|
|
194
|
-
|
|
195
|
-
- The gate decides only whether the user is notified. Reviewer "yes" is input,
|
|
196
|
-
not a veto.
|
|
197
|
-
- `escalate` → records the hold, then one `hermes send` through
|
|
198
|
-
[axstack-relay](../../axstack-relay/SKILL.md) naming the PR, the criterion,
|
|
199
|
-
every reviewer's reason, where the user acts (Orca conversation, worktree, or
|
|
200
|
-
PR), and the hold; it publishes nothing.
|
|
201
|
-
- `proceed` → no notification; a push or `COMMENT` publication then
|
|
202
|
-
additionally requires no unresolved validated blocking finding, because
|
|
203
|
-
`proceed` never overrides a validated blocking finding. A reviewer security
|
|
204
|
-
"yes" that the gate does not escalate is recorded as rejected-with-evidence
|
|
205
|
-
or returned to the author before any mutation.
|
|
206
|
-
- Only `proceed` plus no unresolved validated blocking finding permits a push
|
|
207
|
-
or publication; a push before the gate settles is forbidden.
|
|
208
|
-
- Precedence: credible serious risk found by a reviewer still produces the
|
|
209
|
-
standing internal prompt and dependent-action hold immediately, and the gate
|
|
210
|
-
governs only external notification. That hold is the serious-risk rule of
|
|
211
|
-
[contracts](contracts.md#serious-risk). The internal prompt lands in the run
|
|
212
|
-
record and the automation's Orca conversation; no `hermes send` occurs
|
|
213
|
-
without `escalate`.
|
|
214
|
-
- The health gate takes a watchdog finding and its evidence, not reviewer
|
|
215
|
-
verdicts or a candidate; the same two tokens apply.
|
|
216
|
-
- An unavailable gate or required reviewer records a hold, pauses mutation for
|
|
217
|
-
that PR, and is treated as a watchdog health finding. An unavailable gate
|
|
218
|
-
cannot be escalated through itself: the hold stays, the gap is visible in
|
|
219
|
-
Orca run history and the sidecar, and no substitute or unauthorized send
|
|
220
|
-
occurs.
|
|
221
|
-
|
|
222
|
-
## Precheck (bounded shell, no model)
|
|
223
|
-
|
|
224
|
-
Resolve self; run the three searches with `--json url,number,repository,updatedAt`;
|
|
225
|
-
write the full discovery list to `pending.json`. For each allowlisted,
|
|
226
|
-
non-expired own PR add head SHA, base SHA, and the check rollup of that head via
|
|
227
|
-
`gh pr view --json headRefOid,baseRefOid,statusCheckRollup`; check state is
|
|
228
|
-
data, and only authentication, command, and network errors are `error`. For pair
|
|
229
|
-
C/D only, that call also requests `isDraft` and the precheck adds draft status
|
|
230
|
-
to the hashed fingerprint, so a draft becoming ready wakes the driver on its
|
|
231
|
-
own; pair A/B's queried fields and fingerprint are unchanged. Hash
|
|
232
|
-
only allowlisted PRs. Read `cursor.json` for the last processed fingerprint and
|
|
233
|
-
for due control work: a watch deadline at or before now (pair A/B only, since
|
|
234
|
-
pair C/D stores no deadlines), a pending failed-relay retry, or a per-PR repair
|
|
235
|
-
cap that has expired (pair C/D only, since pair A/B has no caps). Exit 0 when the hash differs or control work is due;
|
|
236
|
-
otherwise exit non-zero. Exit non-zero without running when the previous driver
|
|
237
|
-
run is still active, or on `error`. Append one line
|
|
238
|
-
`<ts> <changed|due|unchanged|error|busy>` to `precheck.log`.
|
|
239
|
-
|
|
240
|
-
Terminal hygiene runs before the searches and covers driver terminals only.
|
|
241
|
-
Ownership is the driver automation's own recorded `terminalPtyId` from its Orca
|
|
242
|
-
run history, never a terminal title: Orca rewrites a Claude terminal's title to
|
|
243
|
-
the agent's current task summary, so a driver's title drifts and an unrelated
|
|
244
|
-
session can acquire one that reads like a driver. The precheck closes the
|
|
245
|
-
previous ticks' idle driver terminals with `--tab` — without it the pane closes
|
|
246
|
-
but the session stays listed and is never reclaimed, which also makes an
|
|
247
|
-
over-match destructive — and a driver terminal that is still working makes the
|
|
248
|
-
tick `busy`. It never closes a watchdog terminal or any terminal outside this
|
|
249
|
-
automation; an unreadable ownership source is `error`, never a silent empty
|
|
250
|
-
sweep. No two of the four automations may share a dispatch minute: A, B, C and D each
|
|
251
|
-
take a distinct minute, so a driver and its watchdog never collide and neither
|
|
252
|
-
pair can disturb the other's terminal hygiene.
|
|
253
|
-
|
|
254
|
-
The observed fingerprint is written to `pending.json`. After the processed
|
|
255
|
-
tick the driver promotes exactly that value to `cursor.json`, never a
|
|
256
|
-
recomputed one, so an event landing during a run is processed on the following
|
|
257
|
-
tick.
|
|
183
|
+
Write the local review file to the workspace review directory
|
|
184
|
+
`~/defi/misc/reviews/` under the existing convention:
|
|
185
|
+
`review-PR-<num>.html` with no prefix means `defi-com/monorepo`;
|
|
186
|
+
`review-mobile-PR-<num>.html` and
|
|
187
|
+
`review-azure-next-hybrid-PR-<num>.html` name the other repositories. A write
|
|
188
|
+
failure is recorded but does not withhold the verdict.
|
|
258
189
|
|
|
259
|
-
|
|
190
|
+
Authored repair commits a local candidate, obtains one Sol review at that
|
|
191
|
+
local SHA and the Luna gate, resolves every validated blocker, records test
|
|
192
|
+
evidence, re-reads remote head/base/draft/deploy set/allowlist, then pushes
|
|
193
|
+
fast-forward. The monorepo unit gate uses `~/.bun-1.2.2/bin/bun`; a suite that
|
|
194
|
+
cannot run locally is an explicit unverified boundary.
|
|
260
195
|
|
|
261
|
-
|
|
262
|
-
identity through Orca runtime inspection and records it; a self-written label
|
|
263
|
-
is not evidence. Expected model: Opus. A mismatch or unknown identity holds
|
|
264
|
-
repair and publication for that run and is a health finding.
|
|
265
|
-
|
|
266
|
-
For each changed PR:
|
|
267
|
-
|
|
268
|
-
- Own PR → [axstack-watch](../../axstack-watch/SKILL.md) on the exact head
|
|
269
|
-
SHA in a per-PR child worktree; the driver worktree never checks out a PR
|
|
270
|
-
branch. Follow its repair-publication reference: candidate committed locally,
|
|
271
|
-
authored review at the local SHA, gate, then fast-forward push with the
|
|
272
|
-
publication readback immediately before it.
|
|
273
|
-
- Peer PR → two isolated `axstack-review` passes on the exact head SHA, then
|
|
274
|
-
the gate, then one owner-synthesized `COMMENT` review under the review skill's
|
|
275
|
-
COMMENT branch. `INCOMPLETE` or an unavailable required reviewer records a
|
|
276
|
-
hold and publishes nothing.
|
|
277
|
-
|
|
278
|
-
Watch window, pair A/B only: 24 hours per own PR from first observation, ending
|
|
279
|
-
early on merge or close. The deadline is stored in `cursor.json`; a due deadline
|
|
280
|
-
wakes the driver through the precheck even when GitHub is unchanged, and the
|
|
281
|
-
driver rechecks the deadline immediately before any publication. Expiry marks
|
|
282
|
-
the PR `expired` in the record and sidecar, records a resumable handoff, and
|
|
283
|
-
stops silently. The driver skips an expired PR until the user re-arms it in the
|
|
284
|
-
automation session, and the precheck ignores it; an expired PR is never silently
|
|
285
|
-
re-adopted. Pair C/D does not use this window at all; see "Watch window" above
|
|
286
|
-
for its rolling replacement, and it stores no deadline and reaches no `expired`
|
|
287
|
-
state.
|
|
288
|
-
|
|
289
|
-
Per-PR event state: head SHA, base SHA, check rollup, and processed request and
|
|
290
|
-
comment IDs. A review receipt is reused only when head, base, and scope are
|
|
291
|
-
unchanged. A new failing check, base change, review request, or qualifying
|
|
292
|
-
comment at an unchanged head is new work. Pair C/D additionally carries draft
|
|
293
|
-
status in that state, and a draft status that has changed from true to false is
|
|
294
|
-
new work for C/D, which is how a draft becoming ready for review reaches its
|
|
295
|
-
driver.
|
|
296
|
-
|
|
297
|
-
## Watchdog tick
|
|
298
|
-
|
|
299
|
-
Before its gate dispatch the watchdog validates its own effective session
|
|
300
|
-
identity through Orca runtime inspection; unknown identity holds the dispatch
|
|
301
|
-
and is itself recorded in `watchdog.json`.
|
|
302
|
-
|
|
303
|
-
Read Orca run history for the driver, `precheck.log`, and the run record.
|
|
304
|
-
Thresholds: three consecutive `error` lines in `precheck.log`; three
|
|
305
|
-
consecutive failed driver runs; no successful driver run within two hours while
|
|
306
|
-
the precheck logged `changed` or `due`; any unrequested fresh-session fallback
|
|
307
|
-
or non-Opus effective identity recorded by the driver; any `failed` relay
|
|
308
|
-
receipt older than one tick or any `uncertain` receipt; a hold with no owner.
|
|
309
|
-
A quiet precheck history with no due work is healthy.
|
|
310
|
-
|
|
311
|
-
The two-hour stall threshold above is pair A/B's. Pair D uses a stall window
|
|
312
|
-
re-derived before enabling as `max(2h, 3 x the 95th-percentile observed C tick
|
|
313
|
-
duration over at least 10 ticks)`, recorded with its sample, because C's
|
|
314
|
-
fingerprint legitimately changes on nearly every tick and a literal two hours
|
|
315
|
-
would fire on the first slow tick. Every other threshold is identical for both
|
|
316
|
-
pairs.
|
|
317
|
-
|
|
318
|
-
Each health finding gets an occurrence id `(type, first-observed UTC
|
|
319
|
-
timestamp)`; it stays deduplicated while unresolved, and a later recurrence is
|
|
320
|
-
a new occurrence. Pass a finding to the gate under the automation-health
|
|
321
|
-
criterion. On `escalate` the watchdog itself performs that one gate-authorized
|
|
322
|
-
`hermes send` and records the receipt in `watchdog.json`; it never mutates
|
|
323
|
-
GitHub and never writes `progress.md` or `cursor.json`. Failed or uncertain
|
|
324
|
-
delivery stays visibly held in `watchdog.json` and Orca run history.
|
|
325
|
-
|
|
326
|
-
## Run record and sidecar
|
|
327
|
-
|
|
328
|
-
One run id per pair for the lifetime of that pair — `20260916-pr-automations`
|
|
329
|
-
for A/B and `20260916-defi-automations` for C/D — each in the
|
|
330
|
-
[run record](run-record.md) shape with that pair's driver as sole writer of its
|
|
331
|
-
own `progress.md`. C/D's run directory lives under the same axstack
|
|
332
|
-
`git-common-dir` as A/B's, in its own `axstack/runs/<run id>/` folder; no
|
|
333
|
-
sidecar, record, or cursor is shared between the pairs. Per PR it stores processed event IDs, exact head and base SHAs, review receipts
|
|
334
|
-
per SHA, gate decisions, `hermes send` receipts with `message_id` and state, and
|
|
335
|
-
holds. The watch deadline and `expired` fields are pair A/B only, since pair C/D
|
|
336
|
-
stores no deadline and cannot reach `expired`; draft status is pair C/D only,
|
|
337
|
-
since only C/D treats a draft transition as new work. Its
|
|
338
|
-
`Notification policy:` line reads, verbatim:
|
|
196
|
+
Every reviewer brief ends exactly:
|
|
339
197
|
|
|
340
198
|
```text
|
|
341
|
-
|
|
199
|
+
Escalate to user: yes | no — <criterion> — <reason>
|
|
342
200
|
```
|
|
343
201
|
|
|
344
|
-
|
|
345
|
-
|
|
346
|
-
|
|
347
|
-
|
|
348
|
-
|
|
349
|
-
|
|
350
|
-
|
|
351
|
-
|
|
352
|
-
|
|
353
|
-
|
|
354
|
-
|
|
355
|
-
`
|
|
356
|
-
|
|
202
|
+
PR work has exactly three criteria: a security concern, a permanent on-chain
|
|
203
|
+
state change, or an architectural change in approach. There is no
|
|
204
|
+
automation-health criterion for reviewers. The Luna gate returns exactly one
|
|
205
|
+
token, `escalate` or `proceed`; there is no gate for health findings.
|
|
206
|
+
`escalate` opens a decision token, sends its message, and exits without waiting
|
|
207
|
+
for a reply. `proceed` does not override a validated blocker. `worker_done`
|
|
208
|
+
names the PR, head, action, GitHub receipt or opened token. The driver alone
|
|
209
|
+
writes the run record.
|
|
210
|
+
|
|
211
|
+
## Decision tokens
|
|
212
|
+
|
|
213
|
+
Store `decisions/<token>.json`, where `<token>` contains at least 96 random
|
|
214
|
+
bits as lowercase hex, produced for example by `openssl rand -hex 16`.
|
|
215
|
+
Immutable bound fields are written once: `repo`, `pr`, `head`, `base`,
|
|
216
|
+
`action` (`approve-verdict`, `request-changes-verdict`, or `push`), plus exact
|
|
217
|
+
`body` and `commit` for verdicts or `candidate_sha`, `head_branch`, and
|
|
218
|
+
`expected_remote_head` for a push. Mutable fields are `state`, `created_at`,
|
|
219
|
+
`decided_at`, `decided_message_id`, `consumed_at`, `receipt`, `send`, and
|
|
220
|
+
`reason`. Never delete a token file.
|
|
221
|
+
|
|
222
|
+
Lifecycle has one named writer per transition, each by temp file + rename:
|
|
223
|
+
|
|
224
|
+
| transition | writer |
|
|
225
|
+
| --- | --- |
|
|
226
|
+
| create `open` | the agent that escalated |
|
|
227
|
+
| `open → approved` / `open → rejected` | the Hermes script only |
|
|
228
|
+
| `approved → spent` / `approved → stale` | the driver only |
|
|
229
|
+
| `rejected → closed` | the driver only |
|
|
230
|
+
| `spent` + `receipt` | the driver only |
|
|
231
|
+
|
|
232
|
+
### Opening
|
|
233
|
+
|
|
234
|
+
Before opening a `push` token, the repair agent pins its candidate with local
|
|
235
|
+
ref `refs/axstack/decisions/<token>` in the project clone, so the candidate
|
|
236
|
+
survives worktree removal. It then sends one
|
|
237
|
+
`hermes send --to telegram` message naming the PR, criterion, every reviewer's
|
|
238
|
+
reason, and the exact replies `approve <token>` and `reject <token>`. Store the
|
|
239
|
+
send receipt. The driver retries a `failed` send once next tick and reconciles
|
|
240
|
+
an `uncertain` send against Hermes output before any retry.
|
|
241
|
+
|
|
242
|
+
The Hermes gateway's fixed `axstack-decide` script accepts only those two exact
|
|
243
|
+
commands. It validates the private user/channel configuration on every call,
|
|
244
|
+
restricts tokens to `^[0-9a-f]{24,}$`, locks and re-reads an `open` token,
|
|
245
|
+
writes only its decision fields by temp file and rename, and prints one line.
|
|
246
|
+
Hermes never runs `gh`, `git`, or `orca`.
|
|
247
|
+
|
|
248
|
+
### Consuming
|
|
249
|
+
|
|
250
|
+
The driver is the only consumer. Immediately before acting it re-reads that
|
|
251
|
+
`state == approved`, then revalidates open/unmerged state, bound head/base,
|
|
252
|
+
absence of a self review for a verdict, or expected remote head plus reachable
|
|
253
|
+
candidate for a push. It writes `spent` before the GitHub call, executes
|
|
254
|
+
exactly the bound action with no second gate, then stores the receipt. A spent
|
|
255
|
+
token without a receipt is
|
|
256
|
+
reconciliation: match the exact review commit/body or destination
|
|
257
|
+
ref/candidate on GitHub; record a match, or the driver retries once under the
|
|
258
|
+
same approval after proving non-execution, or hold ambiguity. A revalidation
|
|
259
|
+
failure writes `stale`
|
|
260
|
+
with the reason and drops the PR's head from `cursor.json`; the driver never
|
|
261
|
+
mints a token. Close rejected files and skip that head. After `spent` or
|
|
262
|
+
`stale`, delete the candidate ref. Tokens do not expire; the watchdog reports
|
|
263
|
+
one open longer than 24 h.
|
|
264
|
+
|
|
265
|
+
## Watchdog
|
|
266
|
+
|
|
267
|
+
The hourly shell precheck only reads `cursor.json`, `precheck.log`,
|
|
268
|
+
`decisions/`, and `orca automations runs --id <driver>`. It performs exactly
|
|
269
|
+
these four checks and always exits non-zero, so no model session launches:
|
|
270
|
+
|
|
271
|
+
| check | trips when |
|
|
272
|
+
| --- | --- |
|
|
273
|
+
| driver stuck | a `changed` or `due` precheck is older than 1 h with no later `tick_done_at`, including a driver that never wrote `tick_started_at` (Orca run status alone is not evidence of completion) |
|
|
274
|
+
| precheck failing | the last three `precheck.log` entries are `error` |
|
|
275
|
+
| worker stuck | a dispatch marker is older than 3 h 30 min and remains uncleared |
|
|
276
|
+
| decision waiting | an `open` decision is older than 24 h |
|
|
277
|
+
|
|
278
|
+
Each trip is `(check, first_observed)`. Write one JSON line per tick to
|
|
279
|
+
`watchdog.log` with `{ts, checks, trips, sent}` and send each occurrence once
|
|
280
|
+
through `hermes send --to telegram`. Persist `sent`, `failed`, or `uncertain`;
|
|
281
|
+
retry `failed` next tick and reconcile `uncertain` before retry. Re-send only
|
|
282
|
+
after the check was observed clear and later recurs. Unreadable evidence is an
|
|
283
|
+
`unknown` occurrence. A healthy tick writes its line and sends nothing. There
|
|
284
|
+
is no gate for health findings.
|
|
285
|
+
|
|
286
|
+
## Run directory
|
|
287
|
+
|
|
288
|
+
One `.git/axstack/runs/<run id>/` directory under the Axstack git-common-dir
|
|
289
|
+
contains:
|
|
290
|
+
|
|
291
|
+
- `cursor.json` — driver only, with these exact keys: `fingerprint`,
|
|
292
|
+
`tick_started_at`, `tick_done_at`, `tick_outcome`, `prs{url: {head, base,
|
|
293
|
+
draft, checks, reviews, last_self_review}}`, `dispatch_markers[]` (`pr`,
|
|
294
|
+
`task_id`, `dispatch_id`, `worktree`, `head`, `started_at`, `reservation`,
|
|
295
|
+
`trigger`), `deferred[]`, `pending_settlement[]`,
|
|
296
|
+
`repair_caps{url: {expires_at}}`, `abandon_count{head: n}`,
|
|
297
|
+
`processed_reviews[]` (`review_id`, `pr`, `head`, `digest`),
|
|
298
|
+
`deploy_on_push{repo: [branches]}`,
|
|
299
|
+
`legacy_automation_reviews[]`, `health[]`;
|
|
300
|
+
- `pending.json`, `precheck.log` — driver precheck only;
|
|
301
|
+
- `decisions/<token>.json` — writers assigned by the lifecycle table;
|
|
302
|
+
- `watchdog.log` and `watchdog-state.json` (occurrence `first_observed` values
|
|
303
|
+
and send receipts) — watchdog only;
|
|
304
|
+
- `progress.md` — driver only, one line per tick plus holds and mention
|
|
305
|
+
readings, with no per-PR prose.
|
|
306
|
+
|
|
307
|
+
Timestamps are UTC `YYYY-MM-DDTHH:MM:SSZ`; an unparsable timestamp is an
|
|
308
|
+
`error` for the precheck and `unknown` for the watchdog, never silently
|
|
309
|
+
ignored.
|
|
310
|
+
|
|
311
|
+
Orca run history is the authoritative log. The launch worktree is a dedicated
|
|
312
|
+
Axstack worktree where nobody develops. The current `defi-automations`
|
|
313
|
+
worktree retires with the old pair. Keep this notification-policy edge in the
|
|
314
|
+
run record: `Notification policy` authorizes the token and watchdog sends;
|
|
315
|
+
delivery uses [axstack-relay](../../axstack-relay/SKILL.md).
|
|
357
316
|
|
|
358
317
|
## Safety holds
|
|
359
318
|
|
|
360
|
-
-
|
|
361
|
-
|
|
362
|
-
|
|
363
|
-
|
|
364
|
-
|
|
365
|
-
|
|
366
|
-
|
|
367
|
-
|
|
368
|
-
|
|
319
|
+
- Non-Opus or unknown driver identity records a hold and dispatches nothing.
|
|
320
|
+
- A GitHub API error makes PR state unknown; never publish or push for it.
|
|
321
|
+
- Enumerate deploy-on-push branches from both repair repositories before
|
|
322
|
+
enabling and store them in `cursor.json`; re-check on allowlist changes.
|
|
323
|
+
- A later tick observing resolution or an explicit user decision clears a
|
|
324
|
+
hold. Silence never clears one.
|
|
325
|
+
|
|
326
|
+
## Cutover
|
|
327
|
+
|
|
328
|
+
Perform this order: the new pair exists disabled; the amended skills and
|
|
329
|
+
references are installed; C and D are disabled; every old driver and worker
|
|
330
|
+
attempt is reconciled to confirmed settlement and each unfinished candidate
|
|
331
|
+
is preserved; the old worktree is removed; the old run directory is made
|
|
332
|
+
read-only; the new pair is enabled at distinct minutes; the first real driver
|
|
333
|
+
tick is recorded. A failure leaves the new pair disabled, and both pairs never
|
|
334
|
+
run together. Historical artefacts are untouched.
|
|
335
|
+
|
|
336
|
+
## Exclusions
|
|
337
|
+
|
|
338
|
+
No obligations table, supersede counter, or review-budget hold. No `COMMENT`
|
|
339
|
+
reviews. No watch deadline or `expired` state. No terminal hygiene or global
|
|
340
|
+
busy guard. No terminal nudge from Hermes. No Hermes access to `gh`, `git`, or
|
|
341
|
+
`orca`. No re-review of human-placed blocks. No per-project state. No changes
|
|
342
|
+
to retired artefacts. No gate for health findings.
|
|
@@ -17,12 +17,9 @@ binding state and receipts to exact revisions.
|
|
|
17
17
|
`axstack-reviewer-secondary` sessions with identical brief and isolated first
|
|
18
18
|
pass; authored = one eligible configured reviewer from actual author
|
|
19
19
|
provenance. Owner and author never review.
|
|
20
|
-
- Driver/monitor/watchdog: the
|
|
21
|
-
|
|
22
|
-
`
|
|
23
|
-
that never sends. `axstack-watchdog` is the hourly health checker that never
|
|
24
|
-
mutates GitHub and performs exactly one kind of send, a gate-authorized
|
|
25
|
-
automation-health escalation recorded in `watchdog.json`.
|
|
20
|
+
- Driver/monitor/watchdog: the driver every 15 minutes dispatches and exits as
|
|
21
|
+
a mutating owner; the watchdog is model-free and read-only, has no gate, and
|
|
22
|
+
records `watchdog.log`; there is no watch deadline for automations.
|
|
26
23
|
- Auditor (`axstack-auditor`): report-only; never edits, merges, activates, or
|
|
27
24
|
audits itself.
|
|
28
25
|
|
|
@@ -99,19 +96,18 @@ Tracking grants no merge, release, model-substitution, or scope authority.
|
|
|
99
96
|
|
|
100
97
|
## Deadline (one rule for every owned timer)
|
|
101
98
|
|
|
102
|
-
The default 24-hour deadline covers
|
|
103
|
-
|
|
104
|
-
|
|
99
|
+
The default 24-hour deadline covers standalone task-owned timers. Stop them at
|
|
100
|
+
deadline and preserve remaining work; there is no watch deadline for
|
|
101
|
+
automations. Merge-ready differs from merged; human merges.
|
|
105
102
|
|
|
106
103
|
## Watch health
|
|
107
104
|
|
|
108
|
-
The user lifted the native-watch hold by user decision on 2026-09-16. The
|
|
109
|
-
automation is a mutating owner for its PRs; `axstack-watchdog` stays
|
|
110
|
-
independent and read-only
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
hold. Build no custom scheduler and use no legacy fallback. Details live in
|
|
105
|
+
The user lifted the native-watch hold by user decision on 2026-09-16. The
|
|
106
|
+
driver automation is a mutating owner for its PRs; `axstack-watchdog` stays
|
|
107
|
+
independent and read-only. The driver every 15 minutes dispatches and exits;
|
|
108
|
+
the watchdog is model-free, has no gate, and records `watchdog.log`; there is
|
|
109
|
+
no watch deadline for automations. Build no custom scheduler and use no legacy
|
|
110
|
+
fallback. Details live in
|
|
115
111
|
[Watch runtime](../../axstack-watch/references/watch-runtime.md).
|
|
116
112
|
|
|
117
113
|
## Audit hook (end of run and meaningful checkpoints)
|
|
@@ -63,12 +63,12 @@ listing all pass.
|
|
|
63
63
|
|
|
64
64
|
## Preserve identity and authority
|
|
65
65
|
|
|
66
|
-
Delivery is one-way. Hermes does not route a
|
|
67
|
-
sending session; its own agent answers replies.
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
66
|
+
Delivery is one-way; no session polls Telegram. Hermes does not route a reply
|
|
67
|
+
back to the sending session; its own agent answers replies. Outside the fixed
|
|
68
|
+
decision-token flow below, a reply is never a receipt, decision, or authority
|
|
69
|
+
for this session, and no persistent owner is needed to send. Every ordinary
|
|
70
|
+
message must say where the user acts: the current Orca conversation, the Orca
|
|
71
|
+
worktree, or the GitHub PR. Do not invent reply commands.
|
|
72
72
|
|
|
73
73
|
Send authority comes from the explicit request or applicable standing policy.
|
|
74
74
|
It grants no merge, publication, ownership-transfer, or model-substitution
|
|
@@ -76,6 +76,24 @@ authority. Delivery is transport evidence only. Revalidate any user decision
|
|
|
76
76
|
that arrives through an authorized channel against the current task and
|
|
77
77
|
existing action boundaries before acting; silence never grants permission.
|
|
78
78
|
|
|
79
|
+
## Decision tokens
|
|
80
|
+
|
|
81
|
+
An automation escalation is the narrow exception defined by
|
|
82
|
+
[Automation sessions](../axstack/references/automations.md). The PR agent opens
|
|
83
|
+
an immutable-bound decision token, sends one message with the exact
|
|
84
|
+
`approve <token>` and `reject <token>` replies, records the send receipt, and
|
|
85
|
+
exits without waiting. The Hermes script decides by validating the private
|
|
86
|
+
channel and updating only an open token file; Hermes never performs the bound
|
|
87
|
+
GitHub or Git action. The driver consumes the file on a later tick, revalidates
|
|
88
|
+
all bound state, marks an approval spent before acting, performs only that
|
|
89
|
+
action, and records its receipt.
|
|
90
|
+
|
|
91
|
+
This does not create a reply channel for the sending agent: Hermes pushes the
|
|
92
|
+
decision to the file through the fixed `axstack-decide` script, and the driver
|
|
93
|
+
reads that file. Delivery is one-way; no session polls Telegram. Token creation,
|
|
94
|
+
decision, and consumption keep their separate writers and authority; a message
|
|
95
|
+
receipt alone authorizes nothing.
|
|
96
|
+
|
|
79
97
|
## Reconcile, deliver, and record
|
|
80
98
|
|
|
81
99
|
Before any new send, check the caller's run record for an existing receipt
|
|
@@ -16,10 +16,10 @@ to select the mode and scope identity, and apply the shared
|
|
|
16
16
|
[PR-shape policy](../axstack/references/pr-shape.md). For an owned candidate,
|
|
17
17
|
load and verify the
|
|
18
18
|
[candidate-publication boundary](../axstack/references/candidate-publication.md).
|
|
19
|
-
When the caller is
|
|
20
|
-
[Automation sessions](../axstack/references/automations.md): its reviewer
|
|
21
|
-
|
|
22
|
-
|
|
19
|
+
When the caller is an Orca driver automation, load
|
|
20
|
+
[Automation sessions](../axstack/references/automations.md): its reviewer briefs
|
|
21
|
+
carry the required escalation field and every eligible peer PR takes a binding
|
|
22
|
+
`APPROVE` or `REQUEST_CHANGES` verdict under the automation exception below.
|
|
23
23
|
|
|
24
24
|
## Peer mode (colleague PR)
|
|
25
25
|
|
|
@@ -219,11 +219,10 @@ Escalate to user: yes | no — <criterion> — <reason>
|
|
|
219
219
|
|
|
220
220
|
Every brief ends with the `Escalate to user` field and the reviewer answers it
|
|
221
221
|
in the receipt. A reviewer may cite only a security concern, a permanent
|
|
222
|
-
on-chain state change, or an architectural change in approach
|
|
223
|
-
|
|
224
|
-
|
|
225
|
-
|
|
226
|
-
for the gate.
|
|
222
|
+
on-chain state change, or an architectural change in approach. Health is not a
|
|
223
|
+
reviewer criterion. The answer is input to the PR escalation gate, not a veto
|
|
224
|
+
and not a verdict; see
|
|
225
|
+
[Automation sessions](../axstack/references/automations.md) for the gate.
|
|
227
226
|
|
|
228
227
|
## Template: review receipt (one block per revision)
|
|
229
228
|
|
|
@@ -258,9 +257,9 @@ private transport values or configuration.
|
|
|
258
257
|
Under an automation session, credible serious risk found by a reviewer still
|
|
259
258
|
raises the standing internal prompt and dependent-action hold immediately, and
|
|
260
259
|
the gate governs only external notification: the internal prompt lands in the
|
|
261
|
-
run record and the automation session's own Orca conversation
|
|
262
|
-
|
|
263
|
-
overrides a validated blocking finding.
|
|
260
|
+
run record and the automation session's own Orca conversation. `escalate`
|
|
261
|
+
opens the bound decision token, sends through `axstack-relay`, and exits
|
|
262
|
+
without waiting; `proceed` never overrides a validated blocking finding.
|
|
264
263
|
|
|
265
264
|
## Publishing rule
|
|
266
265
|
|
|
@@ -306,38 +305,26 @@ for a complete `APPROVE` or `REQUEST_CHANGES` verdict:
|
|
|
306
305
|
Submission is complete only when the remote receipt confirms the review bound
|
|
307
306
|
to the intended commit.
|
|
308
307
|
|
|
309
|
-
Automation exception
|
|
310
|
-
|
|
311
|
-
|
|
312
|
-
|
|
313
|
-
|
|
314
|
-
|
|
315
|
-
|
|
316
|
-
|
|
317
|
-
|
|
318
|
-
|
|
319
|
-
|
|
320
|
-
|
|
321
|
-
|
|
322
|
-
|
|
323
|
-
|
|
324
|
-
|
|
325
|
-
|
|
326
|
-
|
|
327
|
-
|
|
328
|
-
|
|
329
|
-
|
|
330
|
-
|
|
331
|
-
|
|
332
|
-
publication.
|
|
333
|
-
3. Read back the remote head and base immediately before submit; stop if
|
|
334
|
-
either differs from the reviewed candidate.
|
|
335
|
-
4. Submit one owner-synthesized `COMMENT` review bound to the reviewed commit
|
|
336
|
-
through the actual GitHub commit parameter, carrying every receipt's
|
|
337
|
-
findings and limitations, then verify the submission receipt.
|
|
338
|
-
5. A receipt for an unchanged head SHA is never published twice; on ambiguity,
|
|
339
|
-
look up remote state and submit only when absent.
|
|
340
|
-
|
|
341
|
-
Publication is complete when the remote receipt confirms one `COMMENT` review
|
|
342
|
-
bound to the intended commit, or when the recorded hold names the missing
|
|
343
|
-
input and the next owner.
|
|
308
|
+
## Automation exception
|
|
309
|
+
|
|
310
|
+
For a peer PR selected under
|
|
311
|
+
[Automation sessions](../axstack/references/automations.md), apply the same
|
|
312
|
+
complete-review and exact-commit requirements, then run the Luna gate.
|
|
313
|
+
`escalate` opens a bound decision token, sends its message, and exits without
|
|
314
|
+
submitting. `proceed` permits `APPROVE` only with no validated blocker and
|
|
315
|
+
permits `REQUEST_CHANGES` only with at least one evidenced validated blocker.
|
|
316
|
+
`INCOMPLETE`, unavailable inputs, unresolved disagreement, or unknown GitHub
|
|
317
|
+
state submits nothing.
|
|
318
|
+
|
|
319
|
+
Immediately before `gh pr review`, re-read self's reviews at the head. If one
|
|
320
|
+
already exists, skip submission and record its id. Otherwise re-check head,
|
|
321
|
+
base, draft status, authorship, and allowlist, bind the verdict to the exact
|
|
322
|
+
commit parameter, and end the body with this marker line:
|
|
323
|
+
|
|
324
|
+
```text
|
|
325
|
+
<!-- axstack-automation verdict head=<sha> -->
|
|
326
|
+
```
|
|
327
|
+
|
|
328
|
+
After submission, read back and record the review id at the bound head. An
|
|
329
|
+
ambiguous result is looked up before any retry. The automation never submits a
|
|
330
|
+
`COMMENT` review.
|
|
@@ -22,7 +22,7 @@ The authenticated human login selects PRs. Runtime session IDs coordinate work
|
|
|
22
22
|
only and establish neither human identity nor write, reply, or merge authority.
|
|
23
23
|
When the session is the Orca driver automation, also load
|
|
24
24
|
[Automation sessions](../axstack/references/automations.md): it is the owner
|
|
25
|
-
for every PR it handles, and its
|
|
25
|
+
for every PR it handles, and its allowlist bounds every mutation.
|
|
26
26
|
|
|
27
27
|
## 1. Adopt and reconcile
|
|
28
28
|
|
|
@@ -63,9 +63,10 @@ Read-only checks and updates to the already-owned local record need no runtime
|
|
|
63
63
|
load. When the watch needs a new owner or automated observation, first read
|
|
64
64
|
[Watch runtime](references/watch-runtime.md) and then
|
|
65
65
|
[Orca runtime](../axstack/references/orca-runtime.md). Reconcile before creating
|
|
66
|
-
anything. The user lifted the native-watch hold by user decision: the
|
|
67
|
-
|
|
68
|
-
|
|
66
|
+
anything. The user lifted the native-watch hold by user decision: the driver
|
|
67
|
+
every 15 minutes dispatches and exits as a mutating owner; the watchdog is
|
|
68
|
+
model-free and read-only, has no gate, and records `watchdog.log`; there is no
|
|
69
|
+
watch deadline for automations. The driver is the automation session itself,
|
|
69
70
|
with no `axstack-monitor` or `axstack-owner` role row; `axstack-monitor` stays
|
|
70
71
|
an optional read-only observer that never sends. One read-only PR observation
|
|
71
72
|
needs neither.
|
|
@@ -114,10 +115,10 @@ the current revision, and a recorded hold or next owner where work remains.
|
|
|
114
115
|
|
|
115
116
|
When a new actionable event is eligible under a recorded `Notification policy`,
|
|
116
117
|
the owner may use the optional [axstack-relay](../axstack-relay/SKILL.md).
|
|
117
|
-
The monitor never sends
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
118
|
+
The monitor never sends. The automation watchdog never mutates GitHub, has no
|
|
119
|
+
gate, and records occurrences and send receipts in `watchdog.log`; absent
|
|
120
|
+
policy or failed relay uses the current Orca conversation and leaves every
|
|
121
|
+
existing hold open.
|
|
121
122
|
|
|
122
123
|
## 5. State readiness precisely
|
|
123
124
|
|
|
@@ -128,15 +129,13 @@ observed state distinct from merged, and the human merges by default.
|
|
|
128
129
|
|
|
129
130
|
## 6. End and preserve continuity
|
|
130
131
|
|
|
131
|
-
End early when all required PRs merge,
|
|
132
|
-
|
|
133
|
-
|
|
132
|
+
End a standalone watch early when all required PRs merge, at cancellation, or
|
|
133
|
+
at its shared default 24 h deadline. There is no watch deadline for
|
|
134
|
+
automations. In every case, stop and verify all owned registrations.
|
|
134
135
|
|
|
135
136
|
At every end condition, leave the compact state below in the private run record
|
|
136
137
|
and report it in the current chat, even when work remains. Expiry grants neither
|
|
137
|
-
silent renewal nor ownership-transfer authority.
|
|
138
|
-
marks the PR `expired` in the record and sidecar; an `expired` PR is never
|
|
139
|
-
silently re-adopted and is skipped until the user re-arms it.
|
|
138
|
+
silent renewal nor ownership-transfer authority.
|
|
140
139
|
|
|
141
140
|
Transfer ownership through the runtime-owned Orca handoff route only when the
|
|
142
141
|
user explicitly requests it. Before transfer, follow the lifecycle-owned
|
|
@@ -36,9 +36,11 @@ exact new revision and base, all six angles, applicable acceptance, the reply
|
|
|
36
36
|
body identities, and every affected boundary, with no unresolved material
|
|
37
37
|
finding or urgent hold. Under an automation the escalation gate of
|
|
38
38
|
[Automation sessions](../../axstack/references/automations.md) also runs on
|
|
39
|
-
the local SHA
|
|
40
|
-
|
|
41
|
-
|
|
39
|
+
the local SHA. `escalate` creates a token, pins the candidate at
|
|
40
|
+
`refs/axstack/decisions/<token>`, opens the decision token, sends its message,
|
|
41
|
+
and exits instead of pushing. `proceed` permits publication only with no
|
|
42
|
+
unresolved validated blocking finding, and a push before the gate settles is
|
|
43
|
+
forbidden.
|
|
42
44
|
|
|
43
45
|
## 3. Revalidate immediately before publication
|
|
44
46
|
|
|
@@ -4,18 +4,16 @@ Read this before starting, resuming, or stopping automated PR observation.
|
|
|
4
4
|
|
|
5
5
|
## Accepted policy
|
|
6
6
|
|
|
7
|
-
The PR owner remains accountable throughout
|
|
8
|
-
`axstack-monitor` and `axstack-watchdog` are independent,
|
|
9
|
-
authors, reviewers, repliers, or owners.
|
|
7
|
+
The PR owner remains accountable throughout a standalone watch's default
|
|
8
|
+
24-hour window. `axstack-monitor` and `axstack-watchdog` are independent,
|
|
9
|
+
read-only roles, not authors, reviewers, repliers, or owners.
|
|
10
10
|
|
|
11
11
|
- **Monitor:** `axstack-monitor` is an optional read-only observer that reads
|
|
12
|
-
GitHub, all PR feedback, and latest checks
|
|
12
|
+
GitHub, all PR feedback, and latest checks on its registered cadence, persists event
|
|
13
13
|
IDs, wakes the owner only for a new actionable event, and never sends.
|
|
14
|
-
- **Watchdog:**
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
recorded in `watchdog.json`, and otherwise reports a verified health failure
|
|
18
|
-
to the owner.
|
|
14
|
+
- **Watchdog:** for the PR automations, it is a model-free hourly shell precheck
|
|
15
|
+
that never mutates GitHub, has no gate, and records each tick in
|
|
16
|
+
`watchdog.log`.
|
|
19
17
|
|
|
20
18
|
Healthy observations are snapshot-only and update quietly; they wake neither owner nor
|
|
21
19
|
driver. Both roles deduplicate event IDs. Uncertain delivery is reconciled
|
|
@@ -32,25 +30,21 @@ permission pinning are unsupported, and its schedule parser cannot preserve
|
|
|
32
30
|
the accepted bounded expiry by itself. Requested role values or a post-launch
|
|
33
31
|
self-report are not effective launch evidence.
|
|
34
32
|
|
|
35
|
-
The user lifted the native-watch hold by user decision on 2026-09-16. The
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
session itself, with no `axstack-monitor` or `axstack-owner` role row
|
|
40
|
-
materialized for it. The driver records its own model
|
|
41
|
-
identity on every tick and the watchdog compares it with the expected model; a
|
|
42
|
-
mismatch is a safety hold, never a silent substitution. The bounded expiry is
|
|
43
|
-
enforced by the run record's watch deadline, not by the schedule parser. Still
|
|
33
|
+
The user lifted the native-watch hold by user decision on 2026-09-16. The
|
|
34
|
+
driver every 15 minutes dispatches and exits as the mutating owner, and the
|
|
35
|
+
watchdog is model-free and read-only, has no gate, and records `watchdog.log`;
|
|
36
|
+
there is no watch deadline for automations. The driver is the automation
|
|
37
|
+
session itself, with no `axstack-monitor` or `axstack-owner` role row. Still
|
|
44
38
|
introduce no custom scheduler or polling loop and use no legacy runtime
|
|
45
39
|
fallback. The session-level contract lives in
|
|
46
40
|
[Automation sessions](../../axstack/references/automations.md).
|
|
47
41
|
|
|
48
42
|
## Preserve the contract under automation
|
|
49
43
|
|
|
50
|
-
An automation session preserves
|
|
51
|
-
healthy behavior, deduplication, handshake, watched
|
|
52
|
-
|
|
53
|
-
|
|
44
|
+
An automation session preserves its 15-minute driver and hourly model-free
|
|
45
|
+
watchdog cadences, quiet healthy behavior, deduplication, handshake, watched
|
|
46
|
+
scope, and wake owner. Test restart, duplicate ticks, cancellation,
|
|
47
|
+
session-reuse fallback, and final cleanup before enabling.
|
|
54
48
|
A firing timestamp proves neither delivery nor work advancement. An unrequested
|
|
55
49
|
fallback session reconciles ownership and never becomes owner silently.
|
|
56
50
|
|