axstack 0.20.31 → 0.21.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (47) hide show
  1. package/README.md +22 -21
  2. package/bin/axstack.js +17 -5
  3. package/docs/installation.md +97 -48
  4. package/docs/workflows.md +165 -122
  5. package/package.json +3 -3
  6. package/profiles/presets/claude-only.json +23 -23
  7. package/profiles/presets/codex-only.json +10 -10
  8. package/profiles/presets/mixed.json +24 -24
  9. package/skills/axstack/references/automations.md +127 -137
  10. package/skills/axstack/references/autopilot.md +30 -17
  11. package/skills/axstack/references/candidate-publication.md +13 -8
  12. package/skills/axstack/references/contracts.md +10 -8
  13. package/skills/axstack/references/diligence.md +3 -1
  14. package/skills/axstack/references/evidence-archive.md +38 -33
  15. package/skills/axstack/references/lifecycle.md +64 -50
  16. package/skills/axstack/references/review-manager-prompt.md +13 -11
  17. package/skills/axstack/references/role-roster.md +12 -2
  18. package/skills/axstack/references/routing.md +29 -25
  19. package/skills/axstack/references/run-record.md +35 -16
  20. package/skills/axstack/references/t3-runtime.md +234 -0
  21. package/skills/axstack/references/test-audit-weekly.md +62 -0
  22. package/skills/axstack/references/test-value.md +120 -0
  23. package/skills/axstack/references/ui-verification.md +5 -1
  24. package/skills/axstack/references/workspace-hygiene.md +102 -156
  25. package/skills/axstack/scripts/pr-digest.js +120 -0
  26. package/skills/axstack/scripts/resolve-models.js +94 -38
  27. package/skills/axstack-align/SKILL.md +17 -7
  28. package/skills/axstack-audit/SKILL.md +12 -3
  29. package/skills/axstack-audit/references/record.md +1 -1
  30. package/skills/axstack-cleanup/SKILL.md +69 -87
  31. package/skills/axstack-debug/SKILL.md +1 -1
  32. package/skills/axstack-explain/SKILL.md +1 -1
  33. package/skills/axstack-explain/references/visual-qa.md +2 -0
  34. package/skills/axstack-implement/SKILL.md +56 -20
  35. package/skills/axstack-improve/SKILL.md +24 -4
  36. package/skills/axstack-relay/SKILL.md +8 -6
  37. package/skills/axstack-research/SKILL.md +11 -4
  38. package/skills/axstack-review/SKILL.md +34 -30
  39. package/skills/axstack-spec/SKILL.md +18 -13
  40. package/skills/axstack-tickets/SKILL.md +7 -8
  41. package/skills/axstack-watch/SKILL.md +97 -27
  42. package/skills/axstack-watch/references/watch-runtime.md +51 -66
  43. package/src/capabilities.js +33 -69
  44. package/src/installer.js +1 -1
  45. package/src/instructions.js +9 -4
  46. package/skills/axstack/references/orca-runtime.md +0 -202
  47. package/skills/axstack/scripts/trust-path.js +0 -123
@@ -0,0 +1,234 @@
1
+ # T3 runtime boundary
2
+
3
+ Read immediately before role dispatch, receipt consumption, or recovery.
4
+ Axstack owns policy, role selection and evidence; T3 owns threads, runs,
5
+ delegated tasks, messaging and native schedules.
6
+
7
+ ## Preflight and binding
8
+
9
+ The driver must be a T3 thread, save `orchestrator_capabilities` JSON under
10
+ the run record, and follow the advertised tool schema; discovery alone proves
11
+ neither provider readiness nor successful execution. Missing capability holds
12
+ the affected operation, without a substitute runtime.
13
+
14
+ From installed `skills/axstack/roles.json`, the driver must snapshot the selected
15
+ preset and stable role IDs, requested provider/model/class/mode/effort, resolved
16
+ ID, source and time once. Resume preserves that snapshot with no re-resolution;
17
+ changes require the user's explicit decision. Bundled presets are setup inputs.
18
+
19
+ Provider bindings must map codex→`codex`, claude→`claudeAgent`, grok→`grok`,
20
+ antigravity→`antigravity`. Map `modeId` to `runtimeMode:full-access` and effort
21
+ to `options:[{id,value}]`. Stored permission intent is neither effective parity
22
+ nor a security boundary.
23
+
24
+ A preset model must be used as given.
25
+
26
+ `modelClass` must resolve to the newest matching catalog ID for that provider:
27
+ codex `gpt-<N>-<class>`, claude `claude-<class>-<N>-<N>`. Use
28
+ `scripts/resolve-models.js --provider` with the saved capabilities JSON path;
29
+ a missing or malformed catalog holds resolution.
30
+
31
+ A `model:null` role lacking a class must use the first model listed for its
32
+ provider in saved capabilities only for grok and antigravity (launch-by-agent-id
33
+ providers); record the exact ID, rather than an unresolved provider default.
34
+ Antigravity must hold when saved capabilities advertise zero models.
35
+ For codex or claude, a role lacking both model and class is an intentional
36
+ absence and must hold; never use a provider default for that role.
37
+
38
+ An unavailable provider, model, role, mode or effort must hold that role with
39
+ no substitution. Auth, quota, timeout and rejection do not select an alternative.
40
+ Intentional absent seats remain recorded absences; availability is runtime proof.
41
+
42
+ Codex effort must use option ID `reasoningEffort`.
43
+
44
+ Claude effort must use option ID `effort`.
45
+
46
+ Grok effort must use `reasoningEffort` and exclude `max`; its CLI requires ≥1.0.13.
47
+ Advertising Grok alone does not prove that its CLI runs.
48
+
49
+ OpenCode effort must use `variant` (no OpenCode role enters this migration).
50
+
51
+ The dispatch echo must match the requested provider and model; verify options
52
+ and `runtimeMode` by `t3_thread_configuration` read-back. Record requested and
53
+ effective values separately; a worker self-report is not configuration proof.
54
+
55
+ Missing effort read-back must hold roles requiring effort; never record it as
56
+ satisfied. Other providers' effort read-back remains unverified until exercised.
57
+
58
+ The first dispatch for each (provider, model, effort) must reach a running or
59
+ completed run before its siblings launch. Rejection stops siblings of that
60
+ configuration; unrelated configurations remain eligible.
61
+
62
+ ## Role dispatch by permitted writes
63
+
64
+ | Roles | Mechanism and permitted workspace | Completion |
65
+ |---|---|---|
66
+ | Advisers, research, read-only explorers, explainers, diligence, checker, auditor, monitor, arena prose candidates and judges, escalation | Must use async `delegate_task` in the driver worktree, title = dispatch key; tracked and untracked files stay untouched; writes only `<run>/evidence/<key>/` | Native notification followed by persisted `task_status` |
67
+ | Reviewers (peer/authored), release checks, debug investigators, execution investigators (`axstack-explore-execution`), UI verifier (`axstack-ui-verifier`) | Must use async `delegate_task`, title = dispatch key; driver makes a disposable detached checkout of candidate SHA and pinned base with `git worktree add --detach <run>/checkouts/<key> <sha>` (plus pinned debug patch); brief requires `cd` into it; only disposable probes write there, outputs go to `<run>/evidence/<key>/` | Same delegated terminal checks |
68
+ | Author and repairs, code-arena writers | Must use `t3_thread_launch` with `{type:worktree, baseRef:<SHA>, branch:<encoded branch>, startFromOrigin:false}` in their own worktree, kept until PR merges or closes | Writer sends a receipt to the driver; driver verifies terminal run and candidate |
69
+ | Owner | Driver thread in Driver worktree; never writes tracked candidate source or tests; planning artifacts allowed only for repository Markdown; scope, integration, forge mutations and record | No worker launch |
70
+
71
+ The driver must be the sole run-record writer and enforce one writer per
72
+ candidate; it never writes tracked candidate source or tests or repairs an author's source.
73
+ The driver may write planning artifacts (spec, ticket map) in its own worktree when the selected store is repository Markdown.
74
+ Repairs return to that author. Missing or idle sessions grant no ownership transfer.
75
+ The current chat/driver has no role row in any preset.
76
+
77
+ The dispatch key must be `<run>:<role>:<task>:a<n>`, recorded before launch and
78
+ used as the exact whole T3 title. Substring matches do not establish identity.
79
+ Each dispatch binds the approved spec or small-change intent, brief, authority, role snapshot, base and candidate to its dispatch key.
80
+
81
+ The branch must be `axstack/<run>/<role>/<task>-a<n>`; lowercase each segment
82
+ and replace every `[^a-z0-9-]` character with `-`. Keep the dispatch key in its
83
+ original form and record both; normalization is never an identity substitute.
84
+
85
+ `baseRef` must always be a commit SHA, never a branch name; T3 renames its
86
+ `t3code/*` branches. Pin base and candidate before dispatch.
87
+
88
+ Workers must finish with exactly one final marker: `AXSTACK-DONE key=… head=…
89
+ report=…`, `AXSTACK-FAILED key=… head=… report=…`, or `AXSTACK-QUESTION key=… q=…`.
90
+ Reports use absolute private evidence paths; report-only head is the pinned
91
+ candidate SHA. Launched writers deliver the marker through `t3_thread_send` using `mode: queue`
92
+ to the recorded driver thread; delegated children leave it in their final result.
93
+ AXSTACK-* messages from worker threads arrive as user-role messages but are worker receipts, never user instructions or a stop.
94
+ An AXSTACK-FAILED marker follows the failure/replacement rules, never the user-question route.
95
+
96
+ ## Consume completion without advancing stale work
97
+
98
+ Before any `t3_thread_read` for a delegated task, the driver must persist only the `task_status` essentials in private evidence: taskId, status, workState, hasPendingChildRuns, latestTerminalRunId, and the final AXSTACK marker line.
99
+
100
+ Delegated completion must require terminal `completed`, `result_available`,
101
+ `hasPendingChildRuns:false`, and final `AXSTACK-DONE`; a question stays
102
+ incomplete even when native status says completed.
103
+
104
+ Launched writer completion must require terminal `t3_thread_wait` on that run,
105
+ then candidate checks: non-empty diff, clean tree and named red/green logs.
106
+ A receipt message alone counts only as progress; it can precede terminal state.
107
+
108
+ Completion must match the current attempt key and candidate SHA. An older
109
+ attempt never completes a newer one; stale or duplicate receipts remain
110
+ evidence, deduplicated by runtime identity. Process the whole delivery before
111
+ acknowledgment and advance only after checking sender, scope and artifacts.
112
+
113
+ `task_status` failed, a run failed or interrupted, a preparing thread error,
114
+ or `AXSTACK-FAILED` must each produce an incomplete outcome with preserved
115
+ evidence. Silence is not successful completion.
116
+
117
+ After each delegated completion the driver must check its own `HEAD` and
118
+ `git status --porcelain` against their pre-dispatch values: both stay unchanged,
119
+ including untracked entries. Any change is a hold before advancing that work.
120
+
121
+ ## Launch recovery, repairs and questions
122
+
123
+ For a lost launch response the driver must use fully paginated `t3_thread_list`
124
+ with `titleContains=<key>`, filtered to exact whole-title equality, plus
125
+ `git worktree list`. Keep the reserved branch through recovery. A branch or
126
+ worktree without a reconciled thread prevents proof of absence.
127
+
128
+ Recovery must adopt one exact match only after its recorded worktree and branch
129
+ agree; on proven absence relaunch once with the same reserved key and branch.
130
+ Several matches, conflicts, a second uncertain response or incomplete inventory
131
+ hold. Never launch a duplicate writer based on silence.
132
+
133
+ Post-launch the driver must call `t3_thread_wait` with `timeoutMs:120000`:
134
+ failed is launch failure; timed-out with `t3_thread_read` showing both
135
+ `activeRunId` and `worktreePath` is started; still preparing is a hold and
136
+ re-read at the next wake. An unresolved state preserves the attempt.
137
+
138
+ For `AXSTACK-QUESTION` the driver must answer once using `t3_thread_send` to
139
+ the `childThreadId`, record the returned resumed runId, then
140
+ `t3_thread_wait(childThreadId, runId, timeoutMs:600000)`, re-armed by the run
141
+ watch. Persist status again and accept only `latestTerminal*` with
142
+ `latestTerminalRunId` newer than the question run (recorded run ordering, not
143
+ lexical ID order), terminal success and matching receipt key/SHA. The original
144
+ summary stays stable; no notification is assumed for a resumed result.
145
+
146
+ A repair must use `t3_thread_send(writerThreadId, mode:queue)` to the same
147
+ author in the same attempt and worktree; pin the new candidate revision.
148
+
149
+ A replacement after terminal failure must use `a<n+1>`, a new branch and title;
150
+ the failed attempt's branch is kept until salvage. Unknown liveness holds
151
+ replacement; reconcile the old writer before admitting another.
152
+
153
+ With an unsettled launched thread, the driver turn must end only while a bound
154
+ `schedule_task` with `bindToCurrentThread:true`, `everyMs:600000` is armed and
155
+ its ID recorded. Each wake must reconcile all unsettled runs, including a
156
+ writer that died without sending; failed runs hold incomplete work. The watch
157
+ inherits the driver model/workspace and adds no runtime of Axstack's own.
158
+
159
+ Once nothing remains unsettled the driver must `delete_scheduled_task` for the
160
+ run watch and use `list_scheduled_tasks` to read back its absence; an uncertain
161
+ delete preserves the hold and recorded ID.
162
+
163
+ ## Evidence, prompts and authority
164
+
165
+ Put the [Safe-deletion rule](workspace-hygiene.md#safe-deletion) in every worker brief.
166
+ Name the private `<run>/evidence/<key>/` folder in the brief and completion receipt.
167
+
168
+ Before use, commands must scope `TMPDIR` to an owned 0700 directory under the system temp directory, never under `$HOME`, named from the dispatch key and recorded in the receipt.
169
+ Validate its real path, absence of symlinks and ownership before use and cleanup; remove it afterwards by literal absolute path.
170
+ Evidence files still go to the private `<run>/evidence/<key>/` folder.
171
+ Apply equivalent guards to worktree-local paths. Uncertain paths are preserved for reconciliation.
172
+
173
+ Deletion must target an exact validated owned path inside evidence, TMPDIR or
174
+ the worktree, using a literal absolute path or `${VAR:?}`-guarded path: no glob,
175
+ no parent-root deletion; never wipe a general cache. Incidental caches are
176
+ not evidence.
177
+
178
+ The dispatching owner must confirm a worker's own brief question once,
179
+ restating existing authority, then re-verify started state; a second ask holds.
180
+ Never answer trust or permission prompts; brief confirmation adds no authority
181
+ and does not answer a harness or tool dialog.
182
+
183
+ A permission prompt or provider safety refusal must be a held, incomplete
184
+ outcome; never bypass or retry it through another model. Trust, hook review,
185
+ authentication and model prompts also hold; preserve the attempt and inspect
186
+ native state before any authorized recovery.
187
+
188
+ Each reviewer must have a separate checkout and private evidence folder with
189
+ no first-pass cross-read; tracked candidate files remain read-only. Read back
190
+ report and supporting evidence before removing that checkout; evidence already
191
+ outside it needs no archive. Later review gets a fresh checkout; unknown or
192
+ active evidence and unique bytes stay preserved. Untracked files never prove a
193
+ worktree disposable. T3 terminal state alone does not authorize discarding evidence.
194
+
195
+ Input acceptance, started state, effective settings and completed work must
196
+ remain distinct evidence. Silence, contact loss or idle state never proves exit.
197
+ Ordinary resume reconciles the same owner, author, attempt, worktree, revisions
198
+ and pending receipts; uncertainty holds replacement.
199
+
200
+ An explicit user transfer must validate recipient acceptance against intended
201
+ session, scope, revision and authority before changing ownership. The current
202
+ owner remains accountable until then; prior owner stops after acceptance.
203
+ Input acceptance or turn start alone is no transfer receipt.
204
+
205
+ On a T3 `threadId/runId` mismatch the driver must stop consuming and reconcile
206
+ the recorded driver identity with native state; never forge a sender or borrow
207
+ an identity to bypass the mismatch.
208
+
209
+ The driver must retain a user-taken-over T3 thread; never send cleanup commands
210
+ to it, including `t3_thread_organize` settle or archive.
211
+ Require T3 terminal run evidence before `t3_thread_organize` settle or archive;
212
+ these metadata actions are not worktree removal.
213
+
214
+ ## Project preflight and run record
215
+
216
+ `worktreeCleanup` must be `off` for every Axstack project. Preflight reads it
217
+ with `t3_project_read` where exposed; otherwise record a limitation pointing
218
+ to the documented installation setup step. Automatic worktree deletion cannot
219
+ replace evidence readback and salvage.
220
+
221
+ The run record must contain driver threadId, projectId, host, T3 version,
222
+ installed Axstack SHA, capabilities JSON path and scheduledTaskIds for every
223
+ watch and manager schedule.
224
+
225
+ Per dispatch the record must contain key, mechanism, requested target and
226
+ read-back; taskId/childThreadId/childRunId or
227
+ threadId/runId/worktree/branch/base SHA; checkout path with candidate and base
228
+ SHAs; evidence folder, scope/authority, owner, pending receipts, hold and Next.
229
+
230
+ The boundary must provide no Axstack daemon, DB, lock or scheduler. Native T3
231
+ schedules supply wakes; Axstack maintains prose records, not a runtime state
232
+ engine. Private evidence requires explicit publication authority before sharing;
233
+ receipts confer no merge, release, publication, model-substitution, host-mutation
234
+ or expanded scope authority. Human merges by default.
@@ -0,0 +1,62 @@
1
+ # Weekly test-audit prompt
2
+
3
+ You are a fresh finite weekly test-audit session in this repository's dedicated
4
+ T3 project worktree. Before admission, read the activation record: repository,
5
+ test-path allowlist, finite budget, and standing edit and PR-open authority.
6
+ Missing authority, or no passing native canary from [Automations](automations.md),
7
+ holds admission; source checks alone prove no weekly runtime behavior.
8
+ Use T3 `schedule_task` as an unbound weekly `fixed_time` schedule with `bindToCurrentThread:false`.
9
+ T3 `schedule_task` inherits the verified binding read-back and uses a stable
10
+ `clientRequestId`; record its ID, weekly interval and project in the activation record.
11
+ When the finite budget runs out, stop the pass, publish nothing further, and report.
12
+ Use native T3 scheduling and orchestration,
13
+ not a new skill, daemon, scheduler, state engine, or campaign ledger.
14
+
15
+ Reconcile GitHub test-audit PR history and live T3 thread/worktree ownership before
16
+ selecting work. Derive the next single owner boundary within the allowlist from
17
+ the last test-audit PR; with no prior PR, choose the first allowed owner boundary
18
+ and record the choice. Use PR history, never a cursor file or persistent traversal
19
+ state. Never re-propose candidates from any closed unmerged test-audit PR.
20
+ If an open test-audit PR exists, skip the week, publish nothing and report only.
21
+ If paths overlap with live T3 thread or worktree ownership, publish nothing and
22
+ report only.
23
+
24
+ Invoke [Improve](../../axstack-improve/SKILL.md)'s Test-audit lens under [T3 runtime](t3-runtime.md);
25
+ load [Test value](test-value.md), pin the base, and run the baseline suite.
26
+ If a red baseline exists, publish nothing and report only the possible bug.
27
+ If a flaky baseline exists, publish nothing and report only the instability.
28
+ Mark every test declaration in this boundary R/F/C/D and report reviewed and
29
+ eligible counts. Retain uncertain candidates and independent contracts; hidden
30
+ release-only tests are outside scope. There is no deletion quota, score, or target;
31
+ zero deletions is normal.
32
+ If zero proven candidates remain, publish nothing and report only.
33
+
34
+ Route accepted proven C/D to [Implement](../../axstack-implement/SKILL.md) as
35
+ structure-preserving work with the same suite green on pinned base and candidate.
36
+ For every removed assertion, record its location, detectable failure, validation
37
+ command, and named keeper or vacuity/obsolescence proof under Test value.
38
+ For each C and keeper-backed D, prove the keeper red under a targeted disposable
39
+ owner mutation per distinct contract, then restore source byte for byte.
40
+ Improve reports F; the standing authority also permits Implement to repair those
41
+ weak assertions while retaining their contracts.
42
+ For F repairs, apply [F proof](test-value.md#f-proof): base-green,
43
+ removal/inversion-red, byte-for-byte restore, and equivalent-rewording-green.
44
+ Never weaken or loosen an assertion.
45
+ If a weekly PR mixes C/D and F repairs, record both evidence paths in its receipt:
46
+ structure-preserving for C/D and F proof for repairs.
47
+
48
+ The candidate diff touches test files only: deletions, consolidations, F repairs.
49
+ Never edit skip/only/xfail, rewrite snapshots, change coverage-thresholds or CI,
50
+ or edit line-caps. Report test-only production seams without changing them.
51
+ After restoring mutations, verify every non-test path byte-identical to the base;
52
+ coverage, where reported, is a per-file guard and never deletion proof alone.
53
+ Keep one writer and private revision-bound receipts; workers never push.
54
+ Obtain independent review using Implement's configured authored-review roles and
55
+ verify the exact candidate's checks before publication. The driver uses
56
+ `gh stack` and opens at most one test-audit PR per week after independent review;
57
+ the human merges.
58
+
59
+ Notify only under the run's Notification policy: a decision park, merge-ready
60
+ (within the run's milestone cap), or serious-risk hold; never progress or
61
+ heartbeats. Keep routine reports in the T3 driver thread, including skipped or empty passes.
62
+ Settle owned workers and preserve evidence under the shared lifecycle.
@@ -0,0 +1,120 @@
1
+ # Test value
2
+
3
+ Use this reference when authoring or reviewing tests and investigating audit
4
+ candidates. Judge assertions and the production boundary, not test names.
5
+ There is no deletion quota or quality score; zero candidates is valid.
6
+
7
+ ## Authoring gate
8
+
9
+ For every new or changed test, answer:
10
+
11
+ 1. Which observable behavior or independent contract does it protect?
12
+ 2. Which plausible regression would make it fail?
13
+ 3. Why would existing coverage miss that regression? Prefer extending the
14
+ primary boundary test; another layer needs a distinct risk.
15
+ 4. Does it require an export, flag, wrapper, or hook used only by tests?
16
+ If so, exercise the real production boundary instead.
17
+
18
+ Also ask: if every imported function returned `undefined`, would it still pass?
19
+ Investigate a yes. Missing answers or an unproven independent contract fail
20
+ the gate: do not add the test. A test that breaks under a behavior-preserving
21
+ refactor needs a boundary assertion unless it guards an independent contract.
22
+
23
+ For a bug fix, demonstrate the regression test failing on pre-fix code for
24
+ the intended reason, then passing with the fix. Protect the bug once at its
25
+ owner boundary; repeated scenarios at other layers need a separate risk.
26
+
27
+ ## Investigate junk patterns
28
+
29
+ - Probes without assertions; self-comparison; expected values computed by the
30
+ subject being tested.
31
+ - Assertions only about mocks or absence; a mock implementing the behavior
32
+ being asserted; fixtures checking their own supplied results.
33
+ - Repeating constants, config, declared capability flags, or type guarantees
34
+ without independently exercising the promised contract.
35
+ - Exact source, import, or string greps of non-contract text; copied fixtures,
36
+ inventories, manifests, or export lists.
37
+ - Private predicates or call shapes already covered at a boundary; duplicate
38
+ invocations of one contract; tests preserving test-only exports or wrappers.
39
+ - Negative controls passing because an unrelated guard blocked the path;
40
+ names promising behavior the assertions never exercise.
41
+
42
+ The `undefined` check and these patterns are investigative heuristics, never
43
+ automatic deletion verdicts. Read the complete test, production owner and its
44
+ callers, overlapping coverage, CI routing, and relevant history before marking.
45
+
46
+ ## Retention bar
47
+
48
+ Retain independent public API, protocol, config, security, migration, storage,
49
+ platform, package, release, architecture, or prompt-byte contracts. Keep
50
+ observable ordering, credible regression protection, and source inspection
51
+ when it is the cheapest independent guard and survives unrelated rewording or
52
+ identifier changes. Static or slow alone never justifies deletion.
53
+
54
+ A retained test failing on the baseline is a possible product bug: reproduce
55
+ and report it rather than deleting it. Retain uncertain candidates.
56
+
57
+ ## Marks and evidence
58
+
59
+ - **R — retain:** name the independent contract and regression caught.
60
+ - **F — fix assertion:** keep the contract; report the weak assertion for repair.
61
+ - **C — consolidate:** identify the remaining owner test (keeper) first.
62
+ - **D — delete:** identify remaining proof, or prove the assertion vacuous or
63
+ its contract obsolete.
64
+
65
+ For each C/D, record the exact test name and location, what failure it can
66
+ detect, the named keeper or evidence of vacuity/obsolescence, and the validation
67
+ command. Missing evidence leaves the candidate retained and report-only.
68
+
69
+ ## F proof
70
+
71
+ An authorized F repair preserves its contract: the repaired check passes on the
72
+ base, goes red when its instruction or code is removed or inverted by a
73
+ targeted disposable mutation restored byte for byte, and survives equivalent
74
+ rewording for semantic prose. Never weaken or loosen an assertion.
75
+
76
+ ## Deletion proof
77
+
78
+ Pin base and candidate revisions. Run the suite green before and after the
79
+ edit. Map every removed assertion's contract to a remaining keeper or evidence
80
+ that it is vacuous or obsolete; passing suites alone do not prove preservation.
81
+
82
+ For each C and each keeper-backed D, make a targeted disposable mutation of
83
+ the production owner for each distinct contract. Show the named keeper going
84
+ red for that regression, then restore the source byte for byte. Do not infer
85
+ keeper strength from its name or coverage alone.
86
+
87
+ A vacuous D cites the vacuity (no assertion, self-comparison, or an expected
88
+ value derived from the subject). An obsolete D cites the removed production
89
+ path, spec, or history proving the contract is gone. Without that proof,
90
+ report the candidate and retain it. Release-only tests hidden from authoring
91
+ agents are outside weekly audit scope. Where the runner reports coverage,
92
+ use it as a per-file guard; it never authorizes deletion by itself.
93
+
94
+ ## Sources and license
95
+
96
+ Adapted and condensed from OpenClaw's
97
+ [test-audit](https://github.com/openclaw/openclaw/blob/1f351187bd0d/.agents/skills/test-audit/SKILL.md)
98
+ and [campaign](https://github.com/openclaw/openclaw/blob/1f351187bd0d/.agents/skills/test-audit/CAMPAIGN.md).
99
+ MIT License — Copyright (c) 2026 OpenClaw Foundation.
100
+ Paraphrased ideas: [pstack](https://github.com/cursor/plugins/blob/fae2c6ed9582/pstack/skills/principle-test-behavior-not-implementation/SKILL.md)
101
+ (license unverified): test whether assertions depend on subject behavior;
102
+ [Matt Pocock](https://github.com/mattpocock/skills/blob/d81f3a183412/skills/engineering/tdd/tests.md)
103
+ (MIT): use public boundaries and independently chosen expected results.
104
+
105
+ OpenClaw MIT notice:
106
+ Permission is hereby granted, free of charge, to any person obtaining a copy
107
+ of this software and associated documentation files (the "Software"), to deal
108
+ in the Software without restriction, including without limitation the rights
109
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
110
+ copies of the Software, and to permit persons to whom the Software is
111
+ furnished to do so, subject to the following conditions:
112
+ The above copyright notice and this permission notice shall be included in all
113
+ copies or substantial portions of the Software.
114
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
115
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
116
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
117
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
118
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
119
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
120
+ SOFTWARE.
@@ -1,10 +1,14 @@
1
1
  # UI verification
2
2
 
3
3
  Every Playwright, browser, or rendered-UI check, including a "confirm it in the
4
- browser" step, goes through an Orca Dispatch to `axstack-ui-verifier` from the
4
+ browser" step, goes through async `delegate_task` to `axstack-ui-verifier` from the
5
5
  run's role snapshot. Give it the exact build, URL, or artifact and the private
6
6
  dispatch's evidence folder. The verifier is read-only: it never edits source.
7
7
  The PR writer remains the sole writer.
8
+ Read the [T3 runtime boundary](t3-runtime.md) before dispatch and use its
9
+ driver-made disposable detached checkout at the pinned candidate SHA.
10
+ The verifier uses T3 `preview_*` tools for rendered checks.
11
+ Browser and visual checks must run in the delegated `axstack-ui-verifier` in its own detached checkout; outputs go to its private evidence folder, never the driver worktree.
8
12
 
9
13
  Ask for screenshots and observed interactions, accessibility, desktop and
10
14
  mobile layouts, and reduced-motion behavior where relevant. The verifier