axstack 0.20.31 → 0.22.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +25 -23
- package/bin/axstack.js +17 -5
- package/docs/installation.md +104 -51
- package/docs/workflows.md +176 -131
- package/package.json +3 -3
- package/profiles/presets/claude-only.json +23 -23
- package/profiles/presets/codex-only.json +10 -10
- package/profiles/presets/mixed.json +24 -24
- package/skills/axstack/references/automations.md +136 -137
- package/skills/axstack/references/autopilot.md +30 -17
- package/skills/axstack/references/candidate-publication.md +13 -8
- package/skills/axstack/references/contracts.md +13 -12
- package/skills/axstack/references/design-lens.md +3 -3
- package/skills/axstack/references/diligence.md +3 -1
- package/skills/axstack/references/evidence-archive.md +38 -33
- package/skills/axstack/references/lifecycle.md +64 -50
- package/skills/axstack/references/review-manager-prompt.md +13 -11
- package/skills/axstack/references/role-roster.md +12 -2
- package/skills/axstack/references/routing.md +33 -25
- package/skills/axstack/references/run-record.md +35 -16
- package/skills/axstack/references/t3-runtime.md +237 -0
- package/skills/axstack/references/test-audit-weekly.md +62 -0
- package/skills/axstack/references/test-value.md +120 -0
- package/skills/axstack/references/ui-verification.md +5 -1
- package/skills/axstack/references/workspace-hygiene.md +102 -156
- package/skills/axstack/scripts/pr-digest.js +120 -0
- package/skills/axstack/scripts/resolve-models.js +102 -38
- package/skills/axstack-align/SKILL.md +19 -56
- package/skills/axstack-audit/SKILL.md +12 -3
- package/skills/axstack-audit/references/record.md +1 -1
- package/skills/axstack-brainstorm/SKILL.md +24 -0
- package/skills/axstack-brainstorm/references/arena.md +56 -0
- package/skills/axstack-cleanup/SKILL.md +69 -87
- package/skills/axstack-debug/SKILL.md +1 -1
- package/skills/axstack-explain/SKILL.md +1 -1
- package/skills/axstack-explain/references/visual-qa.md +2 -0
- package/skills/axstack-implement/SKILL.md +56 -20
- package/skills/axstack-improve/SKILL.md +24 -4
- package/skills/axstack-relay/SKILL.md +8 -6
- package/skills/axstack-research/SKILL.md +11 -4
- package/skills/axstack-review/SKILL.md +34 -30
- package/skills/axstack-spec/SKILL.md +18 -13
- package/skills/axstack-tickets/SKILL.md +7 -8
- package/skills/axstack-watch/SKILL.md +97 -27
- package/skills/axstack-watch/references/watch-runtime.md +51 -66
- package/src/capabilities.js +33 -69
- package/src/installer.js +1 -1
- package/src/instructions.js +9 -4
- package/skills/axstack/references/orca-runtime.md +0 -202
- package/skills/axstack/scripts/trust-path.js +0 -123
|
@@ -0,0 +1,237 @@
|
|
|
1
|
+
# T3 runtime boundary
|
|
2
|
+
|
|
3
|
+
Read immediately before role dispatch, receipt consumption, or recovery.
|
|
4
|
+
Axstack owns policy, role selection and evidence; T3 owns threads, runs,
|
|
5
|
+
delegated tasks, messaging and native schedules.
|
|
6
|
+
|
|
7
|
+
## Preflight and binding
|
|
8
|
+
|
|
9
|
+
The driver must be a T3 thread, save `orchestrator_capabilities` JSON under
|
|
10
|
+
the run record, and follow the advertised tool schema; discovery alone proves
|
|
11
|
+
neither provider readiness nor successful execution. Missing capability holds
|
|
12
|
+
the affected operation, without a substitute runtime.
|
|
13
|
+
|
|
14
|
+
From installed `skills/axstack/roles.json`, the driver must snapshot the selected
|
|
15
|
+
preset and stable role IDs, requested provider/model/class/mode/effort, resolved
|
|
16
|
+
ID, source and time once. Resume preserves that snapshot with no re-resolution;
|
|
17
|
+
changes require the user's explicit decision. Bundled presets are setup inputs.
|
|
18
|
+
|
|
19
|
+
Provider bindings must map codex→`codex`, claude→`claudeAgent`, grok→`grok`,
|
|
20
|
+
antigravity→`antigravity`. Map `modeId` to `runtimeMode:full-access` and effort
|
|
21
|
+
to `options:[{id,value}]`. Stored permission intent is neither effective parity
|
|
22
|
+
nor a security boundary.
|
|
23
|
+
|
|
24
|
+
A preset model must be used as given.
|
|
25
|
+
|
|
26
|
+
`modelClass` must resolve to the newest matching catalog ID for that provider:
|
|
27
|
+
codex `gpt-<N>-<class>`, claude `claude-<class>-<N>-<N>`. Use
|
|
28
|
+
`scripts/resolve-models.js --provider` with the saved capabilities JSON path;
|
|
29
|
+
a missing or malformed catalog holds resolution.
|
|
30
|
+
|
|
31
|
+
A `model:null` role lacking a class must use the first model listed for its
|
|
32
|
+
provider in saved capabilities only for grok and antigravity (launch-by-agent-id
|
|
33
|
+
providers); record the exact ID, rather than an unresolved provider default.
|
|
34
|
+
Antigravity must hold when saved capabilities advertise zero models.
|
|
35
|
+
For Antigravity, first-listed selection must use the first model ID ending in `-<effort>` because its model ID encodes effort.
|
|
36
|
+
No matching effort suffix holds resolution for Antigravity.
|
|
37
|
+
For Antigravity, pass no effort option; effort read-back uses the model ID suffix.
|
|
38
|
+
For codex or claude, a role lacking both model and class is an intentional
|
|
39
|
+
absence and must hold; never use a provider default for that role.
|
|
40
|
+
|
|
41
|
+
An unavailable provider, model, role, mode or effort must hold that role with
|
|
42
|
+
no substitution. Auth, quota, timeout and rejection do not select an alternative.
|
|
43
|
+
Intentional absent seats remain recorded absences; availability is runtime proof.
|
|
44
|
+
|
|
45
|
+
Codex effort must use option ID `reasoningEffort`.
|
|
46
|
+
|
|
47
|
+
Claude effort must use option ID `effort`.
|
|
48
|
+
|
|
49
|
+
Grok effort must use `reasoningEffort` and exclude `max`; its CLI requires ≥1.0.13.
|
|
50
|
+
Advertising Grok alone does not prove that its CLI runs.
|
|
51
|
+
|
|
52
|
+
OpenCode effort must use `variant` (no OpenCode role enters this migration).
|
|
53
|
+
|
|
54
|
+
The dispatch echo must match the requested provider and model; verify options
|
|
55
|
+
and `runtimeMode` by `t3_thread_configuration` read-back. Record requested and
|
|
56
|
+
effective values separately; a worker self-report is not configuration proof.
|
|
57
|
+
|
|
58
|
+
Missing effort read-back must hold roles requiring effort; never record it as
|
|
59
|
+
satisfied. Other providers' effort read-back remains unverified until exercised.
|
|
60
|
+
|
|
61
|
+
The first dispatch for each (provider, model, effort) must reach a running or
|
|
62
|
+
completed run before its siblings launch. Rejection stops siblings of that
|
|
63
|
+
configuration; unrelated configurations remain eligible.
|
|
64
|
+
|
|
65
|
+
## Role dispatch by permitted writes
|
|
66
|
+
|
|
67
|
+
| Roles | Mechanism and permitted workspace | Completion |
|
|
68
|
+
|---|---|---|
|
|
69
|
+
| Advisers, research, read-only explorers, explainers, diligence, checker, auditor, monitor, arena prose candidates and judges, escalation | Must use async `delegate_task` in the driver worktree, title = dispatch key; tracked and untracked files stay untouched; writes only `<run>/evidence/<key>/` | Native notification followed by persisted `task_status` |
|
|
70
|
+
| Reviewers (peer/authored), release checks, debug investigators, execution investigators (`axstack-explore-execution`), UI verifier (`axstack-ui-verifier`) | Must use async `delegate_task`, title = dispatch key; driver makes a disposable detached checkout of candidate SHA and pinned base with `git worktree add --detach <run>/checkouts/<key> <sha>` (plus pinned debug patch); brief requires `cd` into it; only disposable probes write there, outputs go to `<run>/evidence/<key>/` | Same delegated terminal checks |
|
|
71
|
+
| Author and repairs, code-arena writers | Must use `t3_thread_launch` with `{type:worktree, baseRef:<SHA>, branch:<encoded branch>, startFromOrigin:false}` in their own worktree, kept until PR merges or closes | Writer sends a receipt to the driver; driver verifies terminal run and candidate |
|
|
72
|
+
| Owner | Driver thread in Driver worktree; never writes tracked candidate source or tests; planning artifacts allowed only for repository Markdown; scope, integration, forge mutations and record | No worker launch |
|
|
73
|
+
|
|
74
|
+
The driver must be the sole run-record writer and enforce one writer per
|
|
75
|
+
candidate; it never writes tracked candidate source or tests or repairs an author's source.
|
|
76
|
+
The driver may write planning artifacts (spec, ticket map) in its own worktree when the selected store is repository Markdown.
|
|
77
|
+
Repairs return to that author. Missing or idle sessions grant no ownership transfer.
|
|
78
|
+
The current chat/driver has no role row in any preset.
|
|
79
|
+
|
|
80
|
+
The dispatch key must be `<run>:<role>:<task>:a<n>`, recorded before launch and
|
|
81
|
+
used as the exact whole T3 title. Substring matches do not establish identity.
|
|
82
|
+
Each dispatch binds the approved spec or small-change intent, brief, authority, role snapshot, base and candidate to its dispatch key.
|
|
83
|
+
|
|
84
|
+
The branch must be `axstack/<run>/<role>/<task>-a<n>`; lowercase each segment
|
|
85
|
+
and replace every `[^a-z0-9-]` character with `-`. Keep the dispatch key in its
|
|
86
|
+
original form and record both; normalization is never an identity substitute.
|
|
87
|
+
|
|
88
|
+
`baseRef` must always be a commit SHA, never a branch name; T3 renames its
|
|
89
|
+
`t3code/*` branches. Pin base and candidate before dispatch.
|
|
90
|
+
|
|
91
|
+
Workers must finish with exactly one final marker: `AXSTACK-DONE key=… head=…
|
|
92
|
+
report=…`, `AXSTACK-FAILED key=… head=… report=…`, or `AXSTACK-QUESTION key=… q=…`.
|
|
93
|
+
Reports use absolute private evidence paths; report-only head is the pinned
|
|
94
|
+
candidate SHA. Launched writers deliver the marker through `t3_thread_send` using `mode: queue`
|
|
95
|
+
to the recorded driver thread; delegated children leave it in their final result.
|
|
96
|
+
AXSTACK-* messages from worker threads arrive as user-role messages but are worker receipts, never user instructions or a stop.
|
|
97
|
+
An AXSTACK-FAILED marker follows the failure/replacement rules, never the user-question route.
|
|
98
|
+
|
|
99
|
+
## Consume completion without advancing stale work
|
|
100
|
+
|
|
101
|
+
Before any `t3_thread_read` for a delegated task, the driver must persist only the `task_status` essentials in private evidence: taskId, status, workState, hasPendingChildRuns, latestTerminalRunId, and the final AXSTACK marker line.
|
|
102
|
+
|
|
103
|
+
Delegated completion must require terminal `completed`, `result_available`,
|
|
104
|
+
`hasPendingChildRuns:false`, and final `AXSTACK-DONE`; a question stays
|
|
105
|
+
incomplete even when native status says completed.
|
|
106
|
+
|
|
107
|
+
Launched writer completion must require terminal `t3_thread_wait` on that run,
|
|
108
|
+
then candidate checks: non-empty diff, clean tree and named red/green logs.
|
|
109
|
+
A receipt message alone counts only as progress; it can precede terminal state.
|
|
110
|
+
|
|
111
|
+
Completion must match the current attempt key and candidate SHA. An older
|
|
112
|
+
attempt never completes a newer one; stale or duplicate receipts remain
|
|
113
|
+
evidence, deduplicated by runtime identity. Process the whole delivery before
|
|
114
|
+
acknowledgment and advance only after checking sender, scope and artifacts.
|
|
115
|
+
|
|
116
|
+
`task_status` failed, a run failed or interrupted, a preparing thread error,
|
|
117
|
+
or `AXSTACK-FAILED` must each produce an incomplete outcome with preserved
|
|
118
|
+
evidence. Silence is not successful completion.
|
|
119
|
+
|
|
120
|
+
After each delegated completion the driver must check its own `HEAD` and
|
|
121
|
+
`git status --porcelain` against their pre-dispatch values: both stay unchanged,
|
|
122
|
+
including untracked entries. Any change is a hold before advancing that work.
|
|
123
|
+
|
|
124
|
+
## Launch recovery, repairs and questions
|
|
125
|
+
|
|
126
|
+
For a lost launch response the driver must use fully paginated `t3_thread_list`
|
|
127
|
+
with `titleContains=<key>`, filtered to exact whole-title equality, plus
|
|
128
|
+
`git worktree list`. Keep the reserved branch through recovery. A branch or
|
|
129
|
+
worktree without a reconciled thread prevents proof of absence.
|
|
130
|
+
|
|
131
|
+
Recovery must adopt one exact match only after its recorded worktree and branch
|
|
132
|
+
agree; on proven absence relaunch once with the same reserved key and branch.
|
|
133
|
+
Several matches, conflicts, a second uncertain response or incomplete inventory
|
|
134
|
+
hold. Never launch a duplicate writer based on silence.
|
|
135
|
+
|
|
136
|
+
Post-launch the driver must call `t3_thread_wait` with `timeoutMs:120000`:
|
|
137
|
+
failed is launch failure; timed-out with `t3_thread_read` showing both
|
|
138
|
+
`activeRunId` and `worktreePath` is started; still preparing is a hold and
|
|
139
|
+
re-read at the next wake. An unresolved state preserves the attempt.
|
|
140
|
+
|
|
141
|
+
For `AXSTACK-QUESTION` the driver must answer once using `t3_thread_send` to
|
|
142
|
+
the `childThreadId`, record the returned resumed runId, then
|
|
143
|
+
`t3_thread_wait(childThreadId, runId, timeoutMs:600000)`, re-armed by the run
|
|
144
|
+
watch. Persist status again and accept only `latestTerminal*` with
|
|
145
|
+
`latestTerminalRunId` newer than the question run (recorded run ordering, not
|
|
146
|
+
lexical ID order), terminal success and matching receipt key/SHA. The original
|
|
147
|
+
summary stays stable; no notification is assumed for a resumed result.
|
|
148
|
+
|
|
149
|
+
A repair must use `t3_thread_send(writerThreadId, mode:queue)` to the same
|
|
150
|
+
author in the same attempt and worktree; pin the new candidate revision.
|
|
151
|
+
|
|
152
|
+
A replacement after terminal failure must use `a<n+1>`, a new branch and title;
|
|
153
|
+
the failed attempt's branch is kept until salvage. Unknown liveness holds
|
|
154
|
+
replacement; reconcile the old writer before admitting another.
|
|
155
|
+
|
|
156
|
+
With an unsettled launched thread, the driver turn must end only while a bound
|
|
157
|
+
`schedule_task` with `bindToCurrentThread:true`, `everyMs:600000` is armed and
|
|
158
|
+
its ID recorded. Each wake must reconcile all unsettled runs, including a
|
|
159
|
+
writer that died without sending; failed runs hold incomplete work. The watch
|
|
160
|
+
inherits the driver model/workspace and adds no runtime of Axstack's own.
|
|
161
|
+
|
|
162
|
+
Once nothing remains unsettled the driver must `delete_scheduled_task` for the
|
|
163
|
+
run watch and use `list_scheduled_tasks` to read back its absence; an uncertain
|
|
164
|
+
delete preserves the hold and recorded ID.
|
|
165
|
+
|
|
166
|
+
## Evidence, prompts and authority
|
|
167
|
+
|
|
168
|
+
Put the [Safe-deletion rule](workspace-hygiene.md#safe-deletion) in every worker brief.
|
|
169
|
+
Name the private `<run>/evidence/<key>/` folder in the brief and completion receipt.
|
|
170
|
+
|
|
171
|
+
Before use, commands must scope `TMPDIR` to an owned 0700 directory under the system temp directory, never under `$HOME`, named from the dispatch key and recorded in the receipt.
|
|
172
|
+
Validate its real path, absence of symlinks and ownership before use and cleanup; remove it afterwards by literal absolute path.
|
|
173
|
+
Evidence files still go to the private `<run>/evidence/<key>/` folder.
|
|
174
|
+
Apply equivalent guards to worktree-local paths. Uncertain paths are preserved for reconciliation.
|
|
175
|
+
|
|
176
|
+
Deletion must target an exact validated owned path inside evidence, TMPDIR or
|
|
177
|
+
the worktree, using a literal absolute path or `${VAR:?}`-guarded path: no glob,
|
|
178
|
+
no parent-root deletion; never wipe a general cache. Incidental caches are
|
|
179
|
+
not evidence.
|
|
180
|
+
|
|
181
|
+
The dispatching owner must confirm a worker's own brief question once,
|
|
182
|
+
restating existing authority, then re-verify started state; a second ask holds.
|
|
183
|
+
Never answer trust or permission prompts; brief confirmation adds no authority
|
|
184
|
+
and does not answer a harness or tool dialog.
|
|
185
|
+
|
|
186
|
+
A permission prompt or provider safety refusal must be a held, incomplete
|
|
187
|
+
outcome; never bypass or retry it through another model. Trust, hook review,
|
|
188
|
+
authentication and model prompts also hold; preserve the attempt and inspect
|
|
189
|
+
native state before any authorized recovery.
|
|
190
|
+
|
|
191
|
+
Each reviewer must have a separate checkout and private evidence folder with
|
|
192
|
+
no first-pass cross-read; tracked candidate files remain read-only. Read back
|
|
193
|
+
report and supporting evidence before removing that checkout; evidence already
|
|
194
|
+
outside it needs no archive. Later review gets a fresh checkout; unknown or
|
|
195
|
+
active evidence and unique bytes stay preserved. Untracked files never prove a
|
|
196
|
+
worktree disposable. T3 terminal state alone does not authorize discarding evidence.
|
|
197
|
+
|
|
198
|
+
Input acceptance, started state, effective settings and completed work must
|
|
199
|
+
remain distinct evidence. Silence, contact loss or idle state never proves exit.
|
|
200
|
+
Ordinary resume reconciles the same owner, author, attempt, worktree, revisions
|
|
201
|
+
and pending receipts; uncertainty holds replacement.
|
|
202
|
+
|
|
203
|
+
An explicit user transfer must validate recipient acceptance against intended
|
|
204
|
+
session, scope, revision and authority before changing ownership. The current
|
|
205
|
+
owner remains accountable until then; prior owner stops after acceptance.
|
|
206
|
+
Input acceptance or turn start alone is no transfer receipt.
|
|
207
|
+
|
|
208
|
+
On a T3 `threadId/runId` mismatch the driver must stop consuming and reconcile
|
|
209
|
+
the recorded driver identity with native state; never forge a sender or borrow
|
|
210
|
+
an identity to bypass the mismatch.
|
|
211
|
+
|
|
212
|
+
The driver must retain a user-taken-over T3 thread; never send cleanup commands
|
|
213
|
+
to it, including `t3_thread_organize` settle or archive.
|
|
214
|
+
Require T3 terminal run evidence before `t3_thread_organize` settle or archive;
|
|
215
|
+
these metadata actions are not worktree removal.
|
|
216
|
+
|
|
217
|
+
## Project preflight and run record
|
|
218
|
+
|
|
219
|
+
`worktreeCleanup` must be `off` for every Axstack project. Preflight reads it
|
|
220
|
+
with `t3_project_read` where exposed; otherwise record a limitation pointing
|
|
221
|
+
to the documented installation setup step. Automatic worktree deletion cannot
|
|
222
|
+
replace evidence readback and salvage.
|
|
223
|
+
|
|
224
|
+
The run record must contain driver threadId, projectId, host, T3 version,
|
|
225
|
+
installed Axstack SHA, capabilities JSON path and scheduledTaskIds for every
|
|
226
|
+
watch and manager schedule.
|
|
227
|
+
|
|
228
|
+
Per dispatch the record must contain key, mechanism, requested target and
|
|
229
|
+
read-back; taskId/childThreadId/childRunId or
|
|
230
|
+
threadId/runId/worktree/branch/base SHA; checkout path with candidate and base
|
|
231
|
+
SHAs; evidence folder, scope/authority, owner, pending receipts, hold and Next.
|
|
232
|
+
|
|
233
|
+
The boundary must provide no Axstack daemon, DB, lock or scheduler. Native T3
|
|
234
|
+
schedules supply wakes; Axstack maintains prose records, not a runtime state
|
|
235
|
+
engine. Private evidence requires explicit publication authority before sharing;
|
|
236
|
+
receipts confer no merge, release, publication, model-substitution, host-mutation
|
|
237
|
+
or expanded scope authority. Human merges by default.
|
|
@@ -0,0 +1,62 @@
|
|
|
1
|
+
# Weekly test-audit prompt
|
|
2
|
+
|
|
3
|
+
You are a fresh finite weekly test-audit session in this repository's dedicated
|
|
4
|
+
T3 project worktree. Before admission, read the activation record: repository,
|
|
5
|
+
test-path allowlist, finite budget, and standing edit and PR-open authority.
|
|
6
|
+
Missing authority, or no passing native canary from [Automations](automations.md),
|
|
7
|
+
holds admission; source checks alone prove no weekly runtime behavior.
|
|
8
|
+
Use T3 `schedule_task` as an unbound weekly `fixed_time` schedule with `bindToCurrentThread:false`.
|
|
9
|
+
T3 `schedule_task` inherits the verified binding read-back and uses a stable
|
|
10
|
+
`clientRequestId`; record its ID, weekly interval and project in the activation record.
|
|
11
|
+
When the finite budget runs out, stop the pass, publish nothing further, and report.
|
|
12
|
+
Use native T3 scheduling and orchestration,
|
|
13
|
+
not a new skill, daemon, scheduler, state engine, or campaign ledger.
|
|
14
|
+
|
|
15
|
+
Reconcile GitHub test-audit PR history and live T3 thread/worktree ownership before
|
|
16
|
+
selecting work. Derive the next single owner boundary within the allowlist from
|
|
17
|
+
the last test-audit PR; with no prior PR, choose the first allowed owner boundary
|
|
18
|
+
and record the choice. Use PR history, never a cursor file or persistent traversal
|
|
19
|
+
state. Never re-propose candidates from any closed unmerged test-audit PR.
|
|
20
|
+
If an open test-audit PR exists, skip the week, publish nothing and report only.
|
|
21
|
+
If paths overlap with live T3 thread or worktree ownership, publish nothing and
|
|
22
|
+
report only.
|
|
23
|
+
|
|
24
|
+
Invoke [Improve](../../axstack-improve/SKILL.md)'s Test-audit lens under [T3 runtime](t3-runtime.md);
|
|
25
|
+
load [Test value](test-value.md), pin the base, and run the baseline suite.
|
|
26
|
+
If a red baseline exists, publish nothing and report only the possible bug.
|
|
27
|
+
If a flaky baseline exists, publish nothing and report only the instability.
|
|
28
|
+
Mark every test declaration in this boundary R/F/C/D and report reviewed and
|
|
29
|
+
eligible counts. Retain uncertain candidates and independent contracts; hidden
|
|
30
|
+
release-only tests are outside scope. There is no deletion quota, score, or target;
|
|
31
|
+
zero deletions is normal.
|
|
32
|
+
If zero proven candidates remain, publish nothing and report only.
|
|
33
|
+
|
|
34
|
+
Route accepted proven C/D to [Implement](../../axstack-implement/SKILL.md) as
|
|
35
|
+
structure-preserving work with the same suite green on pinned base and candidate.
|
|
36
|
+
For every removed assertion, record its location, detectable failure, validation
|
|
37
|
+
command, and named keeper or vacuity/obsolescence proof under Test value.
|
|
38
|
+
For each C and keeper-backed D, prove the keeper red under a targeted disposable
|
|
39
|
+
owner mutation per distinct contract, then restore source byte for byte.
|
|
40
|
+
Improve reports F; the standing authority also permits Implement to repair those
|
|
41
|
+
weak assertions while retaining their contracts.
|
|
42
|
+
For F repairs, apply [F proof](test-value.md#f-proof): base-green,
|
|
43
|
+
removal/inversion-red, byte-for-byte restore, and equivalent-rewording-green.
|
|
44
|
+
Never weaken or loosen an assertion.
|
|
45
|
+
If a weekly PR mixes C/D and F repairs, record both evidence paths in its receipt:
|
|
46
|
+
structure-preserving for C/D and F proof for repairs.
|
|
47
|
+
|
|
48
|
+
The candidate diff touches test files only: deletions, consolidations, F repairs.
|
|
49
|
+
Never edit skip/only/xfail, rewrite snapshots, change coverage-thresholds or CI,
|
|
50
|
+
or edit line-caps. Report test-only production seams without changing them.
|
|
51
|
+
After restoring mutations, verify every non-test path byte-identical to the base;
|
|
52
|
+
coverage, where reported, is a per-file guard and never deletion proof alone.
|
|
53
|
+
Keep one writer and private revision-bound receipts; workers never push.
|
|
54
|
+
Obtain independent review using Implement's configured authored-review roles and
|
|
55
|
+
verify the exact candidate's checks before publication. The driver uses
|
|
56
|
+
`gh stack` and opens at most one test-audit PR per week after independent review;
|
|
57
|
+
the human merges.
|
|
58
|
+
|
|
59
|
+
Notify only under the run's Notification policy: a decision park, merge-ready
|
|
60
|
+
(within the run's milestone cap), or serious-risk hold; never progress or
|
|
61
|
+
heartbeats. Keep routine reports in the T3 driver thread, including skipped or empty passes.
|
|
62
|
+
Settle owned workers and preserve evidence under the shared lifecycle.
|
|
@@ -0,0 +1,120 @@
|
|
|
1
|
+
# Test value
|
|
2
|
+
|
|
3
|
+
Use this reference when authoring or reviewing tests and investigating audit
|
|
4
|
+
candidates. Judge assertions and the production boundary, not test names.
|
|
5
|
+
There is no deletion quota or quality score; zero candidates is valid.
|
|
6
|
+
|
|
7
|
+
## Authoring gate
|
|
8
|
+
|
|
9
|
+
For every new or changed test, answer:
|
|
10
|
+
|
|
11
|
+
1. Which observable behavior or independent contract does it protect?
|
|
12
|
+
2. Which plausible regression would make it fail?
|
|
13
|
+
3. Why would existing coverage miss that regression? Prefer extending the
|
|
14
|
+
primary boundary test; another layer needs a distinct risk.
|
|
15
|
+
4. Does it require an export, flag, wrapper, or hook used only by tests?
|
|
16
|
+
If so, exercise the real production boundary instead.
|
|
17
|
+
|
|
18
|
+
Also ask: if every imported function returned `undefined`, would it still pass?
|
|
19
|
+
Investigate a yes. Missing answers or an unproven independent contract fail
|
|
20
|
+
the gate: do not add the test. A test that breaks under a behavior-preserving
|
|
21
|
+
refactor needs a boundary assertion unless it guards an independent contract.
|
|
22
|
+
|
|
23
|
+
For a bug fix, demonstrate the regression test failing on pre-fix code for
|
|
24
|
+
the intended reason, then passing with the fix. Protect the bug once at its
|
|
25
|
+
owner boundary; repeated scenarios at other layers need a separate risk.
|
|
26
|
+
|
|
27
|
+
## Investigate junk patterns
|
|
28
|
+
|
|
29
|
+
- Probes without assertions; self-comparison; expected values computed by the
|
|
30
|
+
subject being tested.
|
|
31
|
+
- Assertions only about mocks or absence; a mock implementing the behavior
|
|
32
|
+
being asserted; fixtures checking their own supplied results.
|
|
33
|
+
- Repeating constants, config, declared capability flags, or type guarantees
|
|
34
|
+
without independently exercising the promised contract.
|
|
35
|
+
- Exact source, import, or string greps of non-contract text; copied fixtures,
|
|
36
|
+
inventories, manifests, or export lists.
|
|
37
|
+
- Private predicates or call shapes already covered at a boundary; duplicate
|
|
38
|
+
invocations of one contract; tests preserving test-only exports or wrappers.
|
|
39
|
+
- Negative controls passing because an unrelated guard blocked the path;
|
|
40
|
+
names promising behavior the assertions never exercise.
|
|
41
|
+
|
|
42
|
+
The `undefined` check and these patterns are investigative heuristics, never
|
|
43
|
+
automatic deletion verdicts. Read the complete test, production owner and its
|
|
44
|
+
callers, overlapping coverage, CI routing, and relevant history before marking.
|
|
45
|
+
|
|
46
|
+
## Retention bar
|
|
47
|
+
|
|
48
|
+
Retain independent public API, protocol, config, security, migration, storage,
|
|
49
|
+
platform, package, release, architecture, or prompt-byte contracts. Keep
|
|
50
|
+
observable ordering, credible regression protection, and source inspection
|
|
51
|
+
when it is the cheapest independent guard and survives unrelated rewording or
|
|
52
|
+
identifier changes. Static or slow alone never justifies deletion.
|
|
53
|
+
|
|
54
|
+
A retained test failing on the baseline is a possible product bug: reproduce
|
|
55
|
+
and report it rather than deleting it. Retain uncertain candidates.
|
|
56
|
+
|
|
57
|
+
## Marks and evidence
|
|
58
|
+
|
|
59
|
+
- **R — retain:** name the independent contract and regression caught.
|
|
60
|
+
- **F — fix assertion:** keep the contract; report the weak assertion for repair.
|
|
61
|
+
- **C — consolidate:** identify the remaining owner test (keeper) first.
|
|
62
|
+
- **D — delete:** identify remaining proof, or prove the assertion vacuous or
|
|
63
|
+
its contract obsolete.
|
|
64
|
+
|
|
65
|
+
For each C/D, record the exact test name and location, what failure it can
|
|
66
|
+
detect, the named keeper or evidence of vacuity/obsolescence, and the validation
|
|
67
|
+
command. Missing evidence leaves the candidate retained and report-only.
|
|
68
|
+
|
|
69
|
+
## F proof
|
|
70
|
+
|
|
71
|
+
An authorized F repair preserves its contract: the repaired check passes on the
|
|
72
|
+
base, goes red when its instruction or code is removed or inverted by a
|
|
73
|
+
targeted disposable mutation restored byte for byte, and survives equivalent
|
|
74
|
+
rewording for semantic prose. Never weaken or loosen an assertion.
|
|
75
|
+
|
|
76
|
+
## Deletion proof
|
|
77
|
+
|
|
78
|
+
Pin base and candidate revisions. Run the suite green before and after the
|
|
79
|
+
edit. Map every removed assertion's contract to a remaining keeper or evidence
|
|
80
|
+
that it is vacuous or obsolete; passing suites alone do not prove preservation.
|
|
81
|
+
|
|
82
|
+
For each C and each keeper-backed D, make a targeted disposable mutation of
|
|
83
|
+
the production owner for each distinct contract. Show the named keeper going
|
|
84
|
+
red for that regression, then restore the source byte for byte. Do not infer
|
|
85
|
+
keeper strength from its name or coverage alone.
|
|
86
|
+
|
|
87
|
+
A vacuous D cites the vacuity (no assertion, self-comparison, or an expected
|
|
88
|
+
value derived from the subject). An obsolete D cites the removed production
|
|
89
|
+
path, spec, or history proving the contract is gone. Without that proof,
|
|
90
|
+
report the candidate and retain it. Release-only tests hidden from authoring
|
|
91
|
+
agents are outside weekly audit scope. Where the runner reports coverage,
|
|
92
|
+
use it as a per-file guard; it never authorizes deletion by itself.
|
|
93
|
+
|
|
94
|
+
## Sources and license
|
|
95
|
+
|
|
96
|
+
Adapted and condensed from OpenClaw's
|
|
97
|
+
[test-audit](https://github.com/openclaw/openclaw/blob/1f351187bd0d/.agents/skills/test-audit/SKILL.md)
|
|
98
|
+
and [campaign](https://github.com/openclaw/openclaw/blob/1f351187bd0d/.agents/skills/test-audit/CAMPAIGN.md).
|
|
99
|
+
MIT License — Copyright (c) 2026 OpenClaw Foundation.
|
|
100
|
+
Paraphrased ideas: [pstack](https://github.com/cursor/plugins/blob/fae2c6ed9582/pstack/skills/principle-test-behavior-not-implementation/SKILL.md)
|
|
101
|
+
(license unverified): test whether assertions depend on subject behavior;
|
|
102
|
+
[Matt Pocock](https://github.com/mattpocock/skills/blob/d81f3a183412/skills/engineering/tdd/tests.md)
|
|
103
|
+
(MIT): use public boundaries and independently chosen expected results.
|
|
104
|
+
|
|
105
|
+
OpenClaw MIT notice:
|
|
106
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
107
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
108
|
+
in the Software without restriction, including without limitation the rights
|
|
109
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
110
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
111
|
+
furnished to do so, subject to the following conditions:
|
|
112
|
+
The above copyright notice and this permission notice shall be included in all
|
|
113
|
+
copies or substantial portions of the Software.
|
|
114
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
115
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
116
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
117
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
118
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
119
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
120
|
+
SOFTWARE.
|
|
@@ -1,10 +1,14 @@
|
|
|
1
1
|
# UI verification
|
|
2
2
|
|
|
3
3
|
Every Playwright, browser, or rendered-UI check, including a "confirm it in the
|
|
4
|
-
browser" step, goes through
|
|
4
|
+
browser" step, goes through async `delegate_task` to `axstack-ui-verifier` from the
|
|
5
5
|
run's role snapshot. Give it the exact build, URL, or artifact and the private
|
|
6
6
|
dispatch's evidence folder. The verifier is read-only: it never edits source.
|
|
7
7
|
The PR writer remains the sole writer.
|
|
8
|
+
Read the [T3 runtime boundary](t3-runtime.md) before dispatch and use its
|
|
9
|
+
driver-made disposable detached checkout at the pinned candidate SHA.
|
|
10
|
+
The verifier uses T3 `preview_*` tools for rendered checks.
|
|
11
|
+
Browser and visual checks must run in the delegated `axstack-ui-verifier` in its own detached checkout; outputs go to its private evidence folder, never the driver worktree.
|
|
8
12
|
|
|
9
13
|
Ask for screenshots and observed interactions, accessibility, desktop and
|
|
10
14
|
mobile layouts, and reduced-motion behavior where relevant. The verifier
|