task-pipeline-skill 1.88.1 → 1.89.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +64 -0
- package/README.md +1 -1
- package/SKILL-CARD.md +1 -1
- package/package.json +3 -3
- package/plugins/task-pipeline/.claude-plugin/plugin.json +1 -1
- package/plugins/task-pipeline/agents/verifier-product.md +3 -1
- package/plugins/task-pipeline/agents/verifier-visual.md +115 -0
- package/plugins/task-pipeline/agents/verifier.md +2 -1
- package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +4 -4
- package/plugins/task-pipeline/skills/task-pipeline/graph.schema.json +22 -1
- package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +4 -4
- package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +12 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/audit.md +14 -8
- package/plugins/task-pipeline/skills/task-pipeline/references/browser.md +97 -3
- package/plugins/task-pipeline/skills/task-pipeline/references/certification.md +46 -5
- package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +28 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/conventions.md +3 -3
- package/plugins/task-pipeline/skills/task-pipeline/references/doctrine-map.md +1 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/grill.md +15 -9
- package/plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md +23 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/portability.md +1 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/spec.md +27 -3
- package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +75 -13
- package/plugins/task-pipeline/skills/task-pipeline/references/work-graph.md +2 -2
- package/plugins/task-pipeline/skills/task-pipeline/scripts/graph.py +45 -16
- package/plugins/task-pipeline/skills/task-pipeline/scripts/stage_checkpoint.py +21 -1
- package/plugins/task-pipeline/skills/task-pipeline/scripts/visual_gate.py +589 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/brief.md +5 -2
- package/plugins/task-pipeline/skills/task-pipeline/templates/browser-claims.json +223 -1
- package/plugins/task-pipeline/skills/task-pipeline/templates/run.md +10 -0
|
@@ -28,6 +28,7 @@ which is how a run says *I checked the browser* and means *I ran the unit tests*
|
|
|
28
28
|
- Sessions, and why an agent needs them
|
|
29
29
|
- Reading a look vs gating on one
|
|
30
30
|
- "Tested in a browser" is three different claims
|
|
31
|
+
- The visual half — pixels against intent, under a contract
|
|
31
32
|
- Getting past a login, and past a backend
|
|
32
33
|
- When the look finds something: debugging the spec that missed it
|
|
33
34
|
- Evidence a reader can open
|
|
@@ -48,7 +49,9 @@ Three consequences the doctrine rests on:
|
|
|
48
49
|
|
|
49
50
|
- **A look costs a page of text and no vision model.** This is why the pipeline can ask
|
|
50
51
|
for one at three stages without the cost being an argument. A `screenshot` exists in
|
|
51
|
-
both channels and *is* pixels — take one for a human
|
|
52
|
+
both channels and *is* pixels — in the functional look, take one for a human, not for
|
|
53
|
+
you to read. Reading pixels against the design's intent is a different check with its
|
|
54
|
+
own contract: *The visual half*, below.
|
|
52
55
|
- **The ref is a fact about the page as rendered**, so `click e12` after a snapshot is
|
|
53
56
|
deterministic in a way a coordinate never is.
|
|
54
57
|
- **A ref that no longer resolves is a finding, not an error to retry past.** The element
|
|
@@ -125,7 +128,7 @@ commonest way a run reports a green it does not have.
|
|
|
125
128
|
|
|
126
129
|
| | What it is | What it proves | Where it counts |
|
|
127
130
|
|---|---|---|---|
|
|
128
|
-
| **The look** | an agent driving a page: open, snapshot, console, network | that this surface renders, right now, and what the browser said while it did | the **look**, stage 6 — recommended, never a gate |
|
|
131
|
+
| **The look** | an agent driving a page: open, snapshot, console, network | that this surface renders, right now, and what the browser said while it did | the **functional look**, stage 6 — recommended, never a gate. Its **visual half** is a gate on a `flagship`, `product` or `ad` surface — *The visual half*, below |
|
|
129
132
|
| **The spec suite** | `playwright test` — the **test runner** | that the assertions someone wrote still hold, on the paths someone thought to write | the **suite** half of the stage-6 gate, counted with every other test |
|
|
130
133
|
| **The library** | `require('playwright')` — `chromium`/`firefox`/`webkit`, `devices`, `request`, `selectors` | whatever your own script asserts; it is an automation API, not a test framework | wherever the project already runs it |
|
|
131
134
|
|
|
@@ -155,6 +158,95 @@ in the claim's own state (the initial screenshot closes nothing about opened/err
|
|
|
155
158
|
a suite PASS closes no look claim; a toggle owes its full cycle; no browser channel
|
|
156
159
|
is NOT_RUN with the reason.
|
|
157
160
|
|
|
161
|
+
## The visual half — pixels against intent, under a contract
|
|
162
|
+
|
|
163
|
+
Everything above reads the **accessibility tree**: it proves the surface renders and the
|
|
164
|
+
browser stayed quiet. It cannot say whether the surface looks like what was designed —
|
|
165
|
+
the wrong weight on a heading, a card that lost its spacing at 200 % text, a dark theme
|
|
166
|
+
that drops a border, a frame that drifted from the Figma it was built from. That is a
|
|
167
|
+
different question, and asking it of a snapshot is how a run reports *it looks right*
|
|
168
|
+
having read no pixels at all. The **functional look** and the **visual half** are two
|
|
169
|
+
checks; neither discharges the other.
|
|
170
|
+
|
|
171
|
+
**When it runs, and when it gates — by the brief's `surface_class`**
|
|
172
|
+
([`stages.md`](stages.md) → stage 0, *The surface class*):
|
|
173
|
+
|
|
174
|
+
| Class | Visual half at stages 5–6 | Human pass at stage 10 |
|
|
175
|
+
|---|---|---|
|
|
176
|
+
| `flagship` | **gate** — full matrix, pairwise across every axis, judge items per the full rubric | **gate** — the approved contact sheet |
|
|
177
|
+
| `product` | **gate** — every state, the mandatory pairs; pairwise holes reported | **gate** — the approved contact sheet |
|
|
178
|
+
| `ad` | **gate** — the ad rubric profile and the safe zones | **gate** — the approved contact sheet |
|
|
179
|
+
| `internal` | recommended — the project linter is the floor; a sheet, if made, is checked for honesty | the functional look closes it |
|
|
180
|
+
|
|
181
|
+
It gates only where the stage-3 VISUAL track ran. A recorded refusal (*«без дизайна»*,
|
|
182
|
+
`Mode: declined`) turns the visual half into the functional look plus the linter, said in
|
|
183
|
+
the close-out — the same rule as every other refusal in the pipeline.
|
|
184
|
+
|
|
185
|
+
**The checks run cheapest first, and the person last:**
|
|
186
|
+
|
|
187
|
+
1. **Deterministic.** The project linter — `python3 scripts/visual_gate.py lint <dir>`,
|
|
188
|
+
which runs sheleg-design's `--lint` where it is installed and answers **NOT_RUN (exit 3)**
|
|
189
|
+
where it is not — plus axe or Lighthouse, the token check, the type checker. An S1
|
|
190
|
+
finding blocks, whatever anyone says about the picture later.
|
|
191
|
+
2. **Regression** against the approved baseline, where one exists (`toHaveScreenshot`, or
|
|
192
|
+
the platform's snapshot test).
|
|
193
|
+
3. **The matrix.** One frame per `SCR-NN` state the screen map lists (default, loading,
|
|
194
|
+
empty, error, offline, long-content, keyboard-up, first-run — the ones that apply),
|
|
195
|
+
across viewport × theme × text size × locale: **pairwise coverage** — every value of one
|
|
196
|
+
axis meets every value of every other in some frame — plus the **mandatory pairs**
|
|
197
|
+
(dark × large text, RTL × narrow). Never the full cross product: every state at every
|
|
198
|
+
combination of every axis value is hundreds of frames, and nobody reads them. Seed the
|
|
199
|
+
worst-case data first (`break-ui`, [`companion-skills.md`](companion-skills.md) →
|
|
200
|
+
*Visual lanes*) so the frames show long names and empty lists, not the demo account.
|
|
201
|
+
4. **Each frame carries its capture record** — revision, route, state, viewport, locale,
|
|
202
|
+
theme, motion, captured-at, source — and is disqualified, not passed, when it is blank,
|
|
203
|
+
stale, of the wrong route or state, or taken before the fonts rendered. This is
|
|
204
|
+
sheleg-design's visual-review contract; the pipeline only refuses a frame without it.
|
|
205
|
+
Where a Figma frame or an approved baseline exists, the frame is **diffed against it**,
|
|
206
|
+
and a failing diff is never a PASS the run writes: either the build drifted, or the
|
|
207
|
+
person approves a new baseline.
|
|
208
|
+
5. **The rubric**, read by a judge that is not the agent that built the surface: a
|
|
209
|
+
checklist per task, never a single score; pairwise only against the approved
|
|
210
|
+
reference and in both orders; three samples, and disagreement is `uncertain`, which
|
|
211
|
+
goes to the person rather than to a coin. **A judge item (J) is `NOT_ASSESSED` until a
|
|
212
|
+
labelled set exists and the judge's agreement with it is measured** — a verdict from an
|
|
213
|
+
uncalibrated judge is a guess with a format. A gate item (G) is deterministic and the
|
|
214
|
+
judge never overrides its FAIL. Every FAIL is a triple: *region → defect → fix*.
|
|
215
|
+
6. **One human pass** over the contact sheet. Approval makes its frames the next baseline.
|
|
216
|
+
|
|
217
|
+
**The contact sheet is the one surface the person reviews**, and its data is
|
|
218
|
+
`templates/browser-claims.json` — the same file, **not a second schema**. A look row
|
|
219
|
+
that carries `axes` is a frame: `axes{viewport,theme,text,locale}`, `capture{revision,
|
|
220
|
+
route,motion,captured_at,source}`, `figma_frame`, `baseline`, `diff`, `rubric[]`. The file
|
|
221
|
+
names its `surface`, `revision`, `review_rounds`, `approved_by` and `approved_at`. The
|
|
222
|
+
viewing page is a self-contained local HTML beside the frames — a full document, `<meta
|
|
223
|
+
charset="utf-8">`, no CDN. **Frames and the HTML stay out of git; the JSON goes in.**
|
|
224
|
+
|
|
225
|
+
```bash
|
|
226
|
+
python3 scripts/visual_gate.py sheet design/review/contact-sheet.json \
|
|
227
|
+
--class product --artifact-root design/review --states SCR-01/default,SCR-01/empty
|
|
228
|
+
# stage 10 adds --require-approval
|
|
229
|
+
```
|
|
230
|
+
|
|
231
|
+
Exit `0` PASS · `1` FAIL · `2` unreadable · `3` NOT_RUN — a frame that did not run makes
|
|
232
|
+
the whole verdict NOT_RUN, never PASS. **On a gated class NOT_RUN stops the stage and
|
|
233
|
+
asks**: connect a capture channel, or the operator accepts the surface as `unverified`,
|
|
234
|
+
recorded in the brief and named in the close-out. It is never reported around, which is
|
|
235
|
+
the failure `DEC-0004` kept the functional look ungated to avoid; `DEC-0006` gates the
|
|
236
|
+
visual half on these classes and keeps that exit open in words.
|
|
237
|
+
|
|
238
|
+
**Two rules keep the loop from becoming the work** ([`loop-guard.md`](loop-guard.md) →
|
|
239
|
+
*The re-render loop*): a re-render budget of **one, two at most**, after which a failing
|
|
240
|
+
item goes to the person as `unresolved`; and **only external, specific feedback** starts a
|
|
241
|
+
round — a linter line, a diff, an audit item, a triple — never "look again". Each return
|
|
242
|
+
from the person adds one to `review_rounds` and one `review:` line to the run ledger, so the
|
|
243
|
+
number of passes a surface took is measured, not remembered.
|
|
244
|
+
|
|
245
|
+
**Native surfaces are not web surfaces.** A web render styled as a phone is a mockup. A
|
|
246
|
+
native screen's frames come from a simulator or a device (XCUITest snapshots, Compose
|
|
247
|
+
screenshot tests) and its accessibility from the platform's own audit; with neither, the
|
|
248
|
+
native rows stay `NOT_RUN`, said in words.
|
|
249
|
+
|
|
158
250
|
## Getting past a login, and past a backend
|
|
159
251
|
|
|
160
252
|
A surface behind auth is the usual reason a run skips the look. Both channels solve it,
|
|
@@ -277,7 +369,9 @@ On a CI box, headed is the failure you will spend an hour on.
|
|
|
277
369
|
| The excuse | Why it fails |
|
|
278
370
|
|---|---|
|
|
279
371
|
| *"`playwright test` is green, the surface is checked."* | The suite asserts what someone wrote down. `DEC-0004`: it is the coverage half, never the look. |
|
|
280
|
-
| *"I took a screenshot, so I looked."* | A screenshot is pixels you did not read. The look is `snapshot` + `console` + `requests`, and the verdict quotes them. |
|
|
372
|
+
| *"I took a screenshot, so I looked."* | A screenshot is pixels you did not read. The functional look is `snapshot` + `console` + `requests`, and the verdict quotes them; reading the pixels is the visual half, with a capture record per frame and a matrix behind it. |
|
|
373
|
+
| *"The snapshot is clean, so it looks right."* | The tree says the heading exists, not that it rendered at the right weight, in the dark theme, at 200 % text. That is the visual half's question, and on a flagship, product or ad surface it is a gate. |
|
|
374
|
+
| *"The judge said it looks great."* | An uncalibrated judge's J items are `NOT_ASSESSED`, and no judge overrides a deterministic FAIL. A verdict needs a checklist, a reference, both orders and three samples. |
|
|
281
375
|
| *"The click failed, I'll find a better selector."* | A ref that stopped resolving **is the finding**. Re-snapshot and report what moved. |
|
|
282
376
|
| *"The docs say the CLI has no `tracing`."* | A vendor page is a claim; `--help` is the tool. This file was written against `--help` **because** a page-derived claim shipped here and was wrong. |
|
|
283
377
|
| *"The tool list is in the docs."* | The page listed tools this version does not ship, and omitted that tracing, video and PDF need `--caps`. Ask the server: 24 tools default, 42 with all caps. |
|
|
@@ -15,6 +15,7 @@ queue is — graph or plan — is what decides, not the mood of the closer.
|
|
|
15
15
|
|
|
16
16
|
- Why one verifier is not enough, stated as the failure it produces
|
|
17
17
|
- The three tiers
|
|
18
|
+
- The fourth reading — `visual`, on a flagship or product surface
|
|
18
19
|
- Blind, and it is the whole design
|
|
19
20
|
- A pass has to mean something, so two rules have teeth
|
|
20
21
|
- The report, and where each field lands
|
|
@@ -60,7 +61,42 @@ that reads no code is not the soft one.
|
|
|
60
61
|
|
|
61
62
|
Agents: [`../../../agents/verifier-unit.md`](../../../agents/verifier-unit.md),
|
|
62
63
|
[`verifier-seam.md`](../../../agents/verifier-seam.md),
|
|
63
|
-
[`verifier-product.md`](../../../agents/verifier-product.md)
|
|
64
|
+
[`verifier-product.md`](../../../agents/verifier-product.md) — and, on a visual surface,
|
|
65
|
+
[`verifier-visual.md`](../../../agents/verifier-visual.md), below.
|
|
66
|
+
|
|
67
|
+
## The fourth reading — `visual`, on a flagship or product surface
|
|
68
|
+
|
|
69
|
+
The three tiers read code, what reaches it, and what the product says about it. **None
|
|
70
|
+
of them opens a picture**, and on a surface whose look is part of the requirement that
|
|
71
|
+
leaves a whole level unread: a node can pass unit, seam and product while its empty
|
|
72
|
+
state renders grey on grey at 200 % text, and every report is truthful about what its
|
|
73
|
+
tier saw.
|
|
74
|
+
|
|
75
|
+
| Tier | Subject | Characteristic finding |
|
|
76
|
+
|---|---|---|
|
|
77
|
+
| `visual` | the contact sheet, the director record, the project linter's output, the rubric | a frame that contradicts the record's intent; a gate item failing under a PASS; a hole in the state × axes matrix; a judge verdict from an uncalibrated judge |
|
|
78
|
+
|
|
79
|
+
**When it runs.** A node that builds a user-facing surface copies the brief's
|
|
80
|
+
`surface_class` ([`stages.md`](stages.md) → stage 0, *The surface class*). On
|
|
81
|
+
**`flagship` and `product`** the `visual` report is **required** — `graph.py certify`
|
|
82
|
+
refuses the round without it and names the tier. On `internal` and `ad` it is accepted
|
|
83
|
+
when given and counted like any other tier, and not demanded: an internal tool's floor
|
|
84
|
+
is the linter, and an ad's gate is its rubric profile on the contact sheet.
|
|
85
|
+
|
|
86
|
+
**It is the fourth blind reading, not a reviewer with a vision model.** Same eight-key
|
|
87
|
+
report, same `breaks`/`risk`, same blindness — it never sees the other three reports
|
|
88
|
+
and they never see it. Its own rules, because a judge of pixels fails in its own ways:
|
|
89
|
+
|
|
90
|
+
- **A checklist per task, never a score.** It answers the rubric's binary items against
|
|
91
|
+
the record and the sheet; "looks polished" is not an item.
|
|
92
|
+
- **Pairwise only against the approved reference, and in both orders**; three samples.
|
|
93
|
+
A verdict that flips with the order, or between samples, is **`uncertain`** and goes to
|
|
94
|
+
the person — it is neither a pass nor a `breaks`.
|
|
95
|
+
- **It never overrides a deterministic FAIL.** A gate item (G) or a linter S1 that
|
|
96
|
+
failed is a `breaks` whatever the picture looks like to it.
|
|
97
|
+
- **A judge item (J) is `NOT_ASSESSED` until a labelled set exists** and the judge's
|
|
98
|
+
agreement with it has been measured. It goes in `not_examined`, which reaches the
|
|
99
|
+
closing verdict as `not_verified` — the honest name for an opinion nobody calibrated.
|
|
64
100
|
|
|
65
101
|
## Blind, and it is the whole design
|
|
66
102
|
|
|
@@ -71,8 +107,10 @@ will paraphrase it back as product truth. The disagreement between blind reading
|
|
|
71
107
|
the instrument, so `graph.py certify` refuses a report whose prose cites another
|
|
72
108
|
tier's verdict.
|
|
73
109
|
|
|
74
|
-
Dispatch all three in one message so they run concurrently
|
|
75
|
-
its `serves`, and the diff —
|
|
110
|
+
Dispatch all three in one message so they run concurrently — all four on a visual
|
|
111
|
+
node. Give each the node id, its `serves`, and the diff — the `visual` tier also the
|
|
112
|
+
paths of the contact sheet, the director record and the linter output — nothing else,
|
|
113
|
+
and never another tier's output.
|
|
76
114
|
|
|
77
115
|
**The second axis, and it is the one an optimisation removes first: whoever produced
|
|
78
116
|
the fix never grades it.** Tier blindness is horizontal — no tier reads another's
|
|
@@ -146,6 +184,9 @@ consumer refuses.
|
|
|
146
184
|
# three reports in, one verdict out — exits 1 if any tier failed
|
|
147
185
|
graph.py certify --node N-007 \
|
|
148
186
|
--tier unit.json --tier seam.json --tier product.json
|
|
187
|
+
# a node with surface_class flagship or product: four, or the round is refused
|
|
188
|
+
graph.py certify --node N-008 \
|
|
189
|
+
--tier unit.json --tier seam.json --tier product.json --tier visual.json
|
|
149
190
|
|
|
150
191
|
# unchanged, and still the only thing that moves the graph
|
|
151
192
|
graph.py close --verdict .task-pipeline/verdict-N-007.json
|
|
@@ -176,8 +217,8 @@ has failed **every** round. A run spinning on one level needs the operator to se
|
|
|
176
217
|
|
|
177
218
|
## What this costs, said out loud
|
|
178
219
|
|
|
179
|
-
Three agents per node instead of one
|
|
180
|
-
paid per node rather than per run. The three are dispatched in parallel, so the
|
|
220
|
+
Three agents per node instead of one — four on a flagship or product surface. That is
|
|
221
|
+
the price of the visibility, and it is paid per node rather than per run. The three are dispatched in parallel, so the
|
|
181
222
|
wall-clock cost is roughly one reading; the token cost is three. A node whose
|
|
182
223
|
`check` is mechanical and whose blast radius is genuinely nil still pays it — and a
|
|
183
224
|
tier with nothing to find says so in `scope` and `not_examined` rather than being
|
|
@@ -24,11 +24,15 @@ one must never look alike.
|
|
|
24
24
|
> stop at the first that answers. The step stays **recommended and never a gate**: a gate
|
|
25
25
|
> an environment cannot satisfy is one an agent learns to report around, and *verified by
|
|
26
26
|
> reading the diff* already prices the absence honestly (`docs/DECISIONS.md`).
|
|
27
|
+
> **`DEC-0006`** scopes that to the functional look: the **visual half** is a gate on a
|
|
28
|
+
> `flagship`, `product` or `ad` surface, and its absent tool reads NOT_RUN and stops to
|
|
29
|
+
> ask rather than passing ([`browser.md`](browser.md) → *The visual half*).
|
|
27
30
|
|
|
28
31
|
## Contents
|
|
29
32
|
|
|
30
33
|
- Built in — nothing to install
|
|
31
34
|
- The matrix
|
|
35
|
+
- Visual lanes — tools, never entry points
|
|
32
36
|
- Optional bridge — substituting an external skill set
|
|
33
37
|
- Preflight (emit before stage 0)
|
|
34
38
|
- Is this skill itself current?
|
|
@@ -76,6 +80,30 @@ one must never look alike.
|
|
|
76
80
|
| ~~superpowers~~ | — | **Not a dependency.** Stages 2/4/5/6 run on the built-in doctrine above. See *Optional bridge* | — |
|
|
77
81
|
| ~~grill-me / grilling~~ | — | **Not a dependency.** The stage-0 grill is built in (`references/grill.md`) | — |
|
|
78
82
|
|
|
83
|
+
## Visual lanes — tools, never entry points
|
|
84
|
+
|
|
85
|
+
The visual half ([`browser.md`](browser.md) → *The visual half*) asks for checks no
|
|
86
|
+
companion above owns alone. These are the public tools that do each one. **Each is a
|
|
87
|
+
tool a stage reaches for, never an entry point and never a second route**: the stage
|
|
88
|
+
decides when, the tool does one job, and the close-out names which one it took. None is
|
|
89
|
+
required; an absent one is a check recorded as NOT_RUN with its reason.
|
|
90
|
+
|
|
91
|
+
| Tool | What it is for | Where in the stages |
|
|
92
|
+
|---|---|---|
|
|
93
|
+
| `break-ui` | seeds worst-case data — the longest name, the empty list, the 4-digit badge, the slow response — so frames show what a real account shows, not the demo | stage 6, **before** the contact-sheet screenshots |
|
|
94
|
+
| `review-animations`, `improve-animations` | reviews motion against its intent: timing, easing, interruption, reduced-motion fallback | stage 6, on a surface with motion; findings as triples |
|
|
95
|
+
| `mobile-native` | the mobile-web surface: viewport, safe areas, touch targets, keyboard-up states | stages 5–6, mobile-web frames of the matrix |
|
|
96
|
+
| `animate-expo` | React Native and Expo motion, built and reviewed on the platform's own primitives | stage 5, an RN or Expo surface |
|
|
97
|
+
| `webapp-testing`, `chrome-devtools` (`take_screenshot`, `lighthouse_audit`) | the frames of the matrix at each viewport and theme, and a Lighthouse pass as the deterministic floor | stage 6, the capture and the floor; stage 8 on a deployed target |
|
|
98
|
+
| `accessibility-review`, `a11y-debugging` | accessibility review of the rendered surface and debugging what it finds | stage 6, beside axe or Lighthouse |
|
|
99
|
+
| XCUITest `performAccessibilityAudit` | the iOS platform's own accessibility audit, run in UI tests | stage 6, a native iOS surface |
|
|
100
|
+
| Compose `enableAccessibilityChecks` | Android's accessibility checks in Compose UI tests | stage 6, a native Android surface |
|
|
101
|
+
| Playwright `toHaveScreenshot` | screenshot regression against the approved baseline, per state and theme | stage 6 (regression), and the nightly or release baseline after |
|
|
102
|
+
| axe-core | the deterministic accessibility floor of a web surface | stage 6, first in the cheap-first order |
|
|
103
|
+
|
|
104
|
+
A native screen is captured on a simulator or a device; a web render styled as a phone is
|
|
105
|
+
a mockup, and the native rows of the sheet stay NOT_RUN without one.
|
|
106
|
+
|
|
79
107
|
## Optional bridge — substituting an external skill set
|
|
80
108
|
|
|
81
109
|
An operator who already runs an equivalent skill set may map it onto stages 2/4/5/6
|
|
@@ -103,9 +103,9 @@ closing a stage with an unread CI verdict.
|
|
|
103
103
|
- Host self-update rules (module docs, runbooks, agent-self cards, etc.) — update
|
|
104
104
|
in the same change. Fix dangling links.
|
|
105
105
|
- **The design destination, on a project with no `docs/ux/`.** When the work uses
|
|
106
|
-
Figma but super-ux isn't in play, there is no `foundation.md` to hold the
|
|
107
|
-
the brief is canonical — and a brief is per-run. Write the team and
|
|
108
|
-
into the host's own docs (`CLAUDE.md`, or the README) in this change, so the next
|
|
106
|
+
Figma but super-ux isn't in play, there is no `foundation.md` to hold the files, so
|
|
107
|
+
the brief is canonical — and a brief is per-run. Write the team and each surface's
|
|
108
|
+
file URL into the host's own docs (`CLAUDE.md`, or the README) in this change, so the next
|
|
109
109
|
run reads the destination instead of creating a second file
|
|
110
110
|
([`grill.md`](grill.md) → *The design destination*).
|
|
111
111
|
- **The code graph:** [graphify](https://github.com/Graphify-Labs/graphify) —
|
|
@@ -30,7 +30,7 @@ UX track on a user-facing task.
|
|
|
30
30
|
| 3 Spec | `references/spec.md` |
|
|
31
31
|
| 4 Plan | `references/planning.md` |
|
|
32
32
|
| the queue the loop walks | `references/work-graph.md` |
|
|
33
|
-
| 5–8 · how a **work-graph node** is CLOSED — three blind readings at three distances, all three required (ceiling 3); a **prose-plan task** closes through `review.md` instead — one reviewer, five-round cap | `references/certification.md` |
|
|
33
|
+
| 5–8 · how a **work-graph node** is CLOSED — three blind readings at three distances, all three required, plus a fourth `visual` reading on a flagship or product surface (ceiling 3); a **prose-plan task** closes through `review.md` instead — one reviewer, five-round cap | `references/certification.md` |
|
|
34
34
|
| 5 Build (worktree, subagents, fix loop) | `references/build.md` + `references/review.md` |
|
|
35
35
|
| 5–6 TDD + suite gate | `references/tdd.md` |
|
|
36
36
|
| 5, 6, 8 The browser — the look, the spec suite, and the difference | `references/browser.md` |
|
|
@@ -18,7 +18,7 @@ coming back to the operator.
|
|
|
18
18
|
- Phase 2 — the gap check, then the loop
|
|
19
19
|
- Domain awareness
|
|
20
20
|
- The autonomy sweep
|
|
21
|
-
- The design destination — one file, decided here, never invented later
|
|
21
|
+
- The design destination — one file per surface, decided here, never invented later
|
|
22
22
|
- The REQ spine — the grill's other hard output
|
|
23
23
|
- Output
|
|
24
24
|
|
|
@@ -170,9 +170,9 @@ explicit "stop and ask me here":
|
|
|
170
170
|
| 0 Docs regime | where settled things live (the decision home — **one** per project, and an existing `docs/adr/` **is** it), who may write it, whether a lease mechanism is present or the run is `ungated`, the gate command and its ratchet floors, and whether this run may raise a floor ([`documentation.md`](documentation.md)) |
|
|
171
171
|
| 1 Docs | external libs/APIs/SDKs in play; any private ones context7 can't resolve → where their docs live |
|
|
172
172
|
| 2 Decompose | is this a platform (several capabilities/surfaces) or one module? if platform: deploy cadence — per module or once at the end |
|
|
173
|
-
| 2–3 Spec | UI verdict (arms super-ux); any scenario-tracing waiver |
|
|
173
|
+
| 2–3 Spec | UI verdict (arms super-ux) **and the surface class** — `flagship`, `product`, `internal` or `ad`, which selects the visual gate profile ([`stages.md`](stages.md) → stage 0, *The surface class*); any scenario-tracing waiver |
|
|
174
174
|
| 3 Design surface | UI tasks only: **Figma on or text-only** (super-ux's project-level choice, default on — check `docs/ux/foundation.md` → *Design tooling* before asking); is the Figma MCP connected; **and if it isn't — ship text-only, or stop here and connect it?** super-ux degrades to text-only on its own and never blocks, which means an unasked question here silently ships a UI feature with no mockups |
|
|
175
|
-
| 3 Design file | Figma on only: **exactly which
|
|
175
|
+
| 3 Design file | Figma on only: **exactly which files — one per surface (App, Web, ASO) — in which team/org** — the recorded ones, or a URL the operator gives, or *create one in a named team* with that creation explicitly authorized. A destination decided at drawing time is how a project ends up with three "design" files and no way to tell which is real. See *The design destination* below |
|
|
176
176
|
| 4–5 Dev | base branch; worktree/branch policy; is `main` off-limits; commit convention; task tracker |
|
|
177
177
|
| 5 Integration | how the branch lands (merge / PR + approver / "leave it unmerged"); parallel fan-out wanted (one worktree per implementer)? |
|
|
178
178
|
| 6 Tests | the test command; what "green" means here; known-red baseline; coverage expectation |
|
|
@@ -188,7 +188,7 @@ preconditions ("staging once lint and the full suite are green; production alway
|
|
|
188
188
|
asks"). Specific and recorded → it satisfies the stage-7 manual gate. Broader,
|
|
189
189
|
absent or ambiguous → stage 7 stops and asks.
|
|
190
190
|
|
|
191
|
-
## The design destination — one file, decided here, never invented later
|
|
191
|
+
## The design destination — one file per surface, decided here, never invented later
|
|
192
192
|
|
|
193
193
|
When the project designs in Figma, **the destination is a stage-0 decision, not a
|
|
194
194
|
stage-3 side effect.** Left to drawing time, the question "where do I put this?"
|
|
@@ -199,16 +199,22 @@ the team actually opens.
|
|
|
199
199
|
|
|
200
200
|
**Settle three things, in this order:**
|
|
201
201
|
|
|
202
|
-
1. **Is there already a file?** Read `docs/ux/foundation.md` →
|
|
203
|
-
first. A recorded, resolving file ends the question — record "use
|
|
204
|
-
file" and move on. Do not ask the operator something the project
|
|
202
|
+
1. **Is there already a file for this surface?** Read `docs/ux/foundation.md` →
|
|
203
|
+
*Design tooling* first. A recorded, resolving file ends the question — record "use
|
|
204
|
+
the recorded file" and move on. Do not ask the operator something the project
|
|
205
|
+
already answered.
|
|
205
206
|
2. **Which team / organization**, by name. A file URL identifies a file; it does not
|
|
206
207
|
say whose workspace it lives in, and a design that lands in someone's personal
|
|
207
208
|
drafts instead of the team space is invisible to everyone who needs it. When the
|
|
208
209
|
operator belongs to several teams, the choice is theirs and it gets written down —
|
|
209
210
|
`whoami` tells you which are available, it does not tell you which is right.
|
|
210
|
-
3. **Which file** — an existing URL the operator supplies, or **creation
|
|
211
|
-
named team, explicitly authorized.**
|
|
211
|
+
3. **Which file, per surface** — an existing URL the operator supplies, or **creation
|
|
212
|
+
in that named team, explicitly authorized.** The shape is **one file per surface** —
|
|
213
|
+
App, Web, ASO (store screenshots, icon, logo) — never one file holding every frame:
|
|
214
|
+
the surfaces have different owners, sizes and reviewers, and a single file is the
|
|
215
|
+
one nobody can hand to any of them. The record names each file with its surface,
|
|
216
|
+
and stage 3's gate checks every frame link against that set
|
|
217
|
+
(`scripts/visual_gate.py filekeys`).
|
|
212
218
|
|
|
213
219
|
**Creating a file in a shared workspace is outward and irreversible enough to need a
|
|
214
220
|
named target.** It follows the same floor as deploy authorization above: *"create
|
|
@@ -23,6 +23,7 @@ searches.
|
|
|
23
23
|
- Bookkeeping — the thing that makes detection mechanical
|
|
24
24
|
- Detection — any one of these trips the guard
|
|
25
25
|
- The review loop — a cap that measures rather than stops
|
|
26
|
+
- The re-render loop — a budget of one, two at most
|
|
26
27
|
- The break protocol
|
|
27
28
|
- When to stop and hand back
|
|
28
29
|
- Rationalizations
|
|
@@ -110,6 +111,28 @@ that was not.
|
|
|
110
111
|
`touch:` lines at the review stage. A round that finds nothing ends the loop by
|
|
111
112
|
definition and needs no counting.
|
|
112
113
|
|
|
114
|
+
## The re-render loop — a budget of one, two at most
|
|
115
|
+
|
|
116
|
+
A visual surface has a loop of its own: render, critique the render, render again. It
|
|
117
|
+
is the one loop here with a **budget rather than a cap**, because its gain is measured
|
|
118
|
+
to flatten fast — past the first or second refinement the change sits inside the noise
|
|
119
|
+
of the judge reading it, and a third round mostly swaps one defect for another. So:
|
|
120
|
+
|
|
121
|
+
- **One re-render per critique, two at the most.** After the second, the item still
|
|
122
|
+
failing is marked **`unresolved`** on the contact sheet with its triple and goes to the
|
|
123
|
+
person on their one pass — not into round three. Keep every rendered version; the last
|
|
124
|
+
is not automatically the best, and the sheet can show two side by side.
|
|
125
|
+
- **Only external, specific feedback starts a round**: a linter finding, a failed audit
|
|
126
|
+
item, a diff against the frame or the approved baseline, a *region → defect → fix*
|
|
127
|
+
triple. "Look again" or "make it better" is not feedback and starts nothing — a round
|
|
128
|
+
with no named defect is churn with a screenshot.
|
|
129
|
+
- **The person's rounds are counted, not remembered.** Each return of the contact sheet
|
|
130
|
+
with triples is a `review:` line in the run ledger ([`../templates/run.md`](../templates/run.md)),
|
|
131
|
+
and the sheet's `review_rounds` carries the same number
|
|
132
|
+
([`browser.md`](browser.md) → *The visual half*). Like the review cap above it is a
|
|
133
|
+
measurement — a class of defect that keeps reaching the person is a rule missing from
|
|
134
|
+
the machine checks, and the fix is that rule, not more attention.
|
|
135
|
+
|
|
113
136
|
## The break protocol
|
|
114
137
|
|
|
115
138
|
When the guard trips, **stop editing immediately**. Do not dispatch another fix, do
|
|
@@ -74,7 +74,7 @@ a row pointing outside the bundle is the defect this file exists to catch.
|
|
|
74
74
|
| What a spec must lock, the UX-track order, the module dossier | `references/spec.md` |
|
|
75
75
|
| The zero-context plan format, parallel groups, set equality | `references/planning.md` |
|
|
76
76
|
| The work graph: its fields, the verbs and their exit codes, and the three invariants a schema cannot state | `references/work-graph.md`, `scripts/graph.py`, `graph.schema.json` |
|
|
77
|
-
| How a node is CLOSED: three blind readings at escalating visibility, all three required
|
|
77
|
+
| How a node is CLOSED: three blind readings at escalating visibility, all three required — four on a flagship or product surface — and the round ledger the ceiling reads | `references/certification.md`, `agents/verifier-{unit,seam,product,visual}.md`, `scripts/graph.py certify` |
|
|
78
78
|
| Workspace isolation, the subagent loop, who may write the register | `references/build.md` |
|
|
79
79
|
| The review rubric, diff packages, the three verdicts | `references/review.md` |
|
|
80
80
|
| **False success** — the class, its known shapes and its two rules | `references/gates.md` |
|
|
@@ -33,9 +33,9 @@ Runs on **super-ux** — the one companion this pipeline recommends by name
|
|
|
33
33
|
give the install line and stop; don't improvise a half-chain.
|
|
34
34
|
|
|
35
35
|
0. **The design destination is already decided — read it, don't re-open it.** When
|
|
36
|
-
Figma is on, the stage-0 brief names the team/org and the
|
|
37
|
-
([`grill.md`](grill.md) → *The design destination*), and
|
|
38
|
-
`docs/ux/foundation.md` → *Design tooling* is the canonical record. Confirm
|
|
36
|
+
Figma is on, the stage-0 brief names the team/org and the files — one per surface
|
|
37
|
+
(App, Web, ASO) — ([`grill.md`](grill.md) → *The design destination*), and
|
|
38
|
+
`docs/ux/foundation.md` → *Design tooling* is the canonical record. Confirm each
|
|
39
39
|
recorded file **resolves** before any drawing. **Never create a file when a
|
|
40
40
|
recorded one resolves; if it doesn't resolve, stop and ask — never create a
|
|
41
41
|
replacement.** A creation happens at most once per project, in the team the
|
|
@@ -90,6 +90,30 @@ green over both.
|
|
|
90
90
|
boundary with Figma (tokens as variables, never raw values carried across). Not
|
|
91
91
|
through it: a purely structural change — what sits where is the UX track's — text,
|
|
92
92
|
a backend, an internal script.
|
|
93
|
+
- **The VISUAL track leaves a trace, and the gate reads the trace.** "The track ran" is
|
|
94
|
+
not checkable; a record is. The track writes a **director record** in the product's
|
|
95
|
+
repository, beside `docs/ux/`: `docs/design/<surface>/director-record.md`, a header
|
|
96
|
+
line `surface_class: <class>` and one `## <Field>` heading per field. Which fields are
|
|
97
|
+
owed is the brief's `surface_class` ([`stages.md`](stages.md) → stage 0, *The surface
|
|
98
|
+
class*):
|
|
99
|
+
|
|
100
|
+
| Class | Fields the record owes |
|
|
101
|
+
|---|---|
|
|
102
|
+
| `flagship` | Brief, Mode, Taste, References, Cast, Fork, Rubric, Critique, Markers, Alignment, Quality, Signature, Surfaces, Haptics, ADA, Open |
|
|
103
|
+
| `product` | Brief, Mode, References, Markers, Open |
|
|
104
|
+
| `ad` | Brief, Mode, References, Markers, ADA (the ad rubric profile and safe zones), Open |
|
|
105
|
+
| `internal` | none — the project linter is the floor |
|
|
106
|
+
|
|
107
|
+
What each field must say is sheleg-design's contract, and its validator checks it:
|
|
108
|
+
`npx sheleg-design-skill --check-record <file>`. The pipeline's gate is `python3
|
|
109
|
+
scripts/visual_gate.py record <file> --class <surface_class>`: it checks the headings
|
|
110
|
+
itself — present, and not a placeholder — and runs that validator where sheleg-design
|
|
111
|
+
is installed and new enough. Where it is not, the validator reads **NOT_RUN**, never
|
|
112
|
+
PASS, and the gate stands on the heading floor with that said. **The refusal is the
|
|
113
|
+
same file**: `## Mode` saying `declined` and why — *«без дизайна»*, *as is* — passes,
|
|
114
|
+
and a bare `declined` with no reason does not. The Rubric is written **before** any
|
|
115
|
+
direction is rendered; a rubric written after the render grades the render it already
|
|
116
|
+
liked.
|
|
93
117
|
- **Each track's refusal is a sentence, never a silence.** *"Без дизайна" / "as is"*
|
|
94
118
|
ends the visual track; *"без бренда" / "draft"* ends the copy track. Either one is
|
|
95
119
|
the operator's to make and costs nothing — but it is **recorded in the brief and
|
|
@@ -187,10 +187,28 @@ never that the work was skipped quietly.
|
|
|
187
187
|
- **UI early-detect:** one branch of the grill is always "does this touch a
|
|
188
188
|
user-facing surface (web/mobile/CLI/TUI)?". If yes → surface **super-ux**
|
|
189
189
|
now (use it if installed; otherwise give the install line — see SKILL.md
|
|
190
|
-
*Prerequisites*); this arms the stage-3 UX track.
|
|
190
|
+
*Prerequisites*); this arms the stage-3 UX track. **And the same branch records
|
|
191
|
+
the surface's class** — the next bullet.
|
|
192
|
+
- **The surface class.** Every user-facing task's brief carries
|
|
193
|
+
`surface_class: flagship | product | internal | ad`, and the class selects the gate
|
|
194
|
+
profile the visual layer is held to for the rest of the run:
|
|
195
|
+
|
|
196
|
+
| Class | What it is | Director record (stage 3) | Visual half (stages 5–6) | Stage 10 |
|
|
197
|
+
|---|---|---|---|---|
|
|
198
|
+
| `flagship` | the surface a product is judged by — landing, onboarding, paywall, a hero screen | the full record | **gate**: full matrix, pairwise across every axis, the full rubric | approved contact sheet |
|
|
199
|
+
| `product` | an ordinary screen of the product | the short record — Brief, Mode, References, Markers, Open | **gate**: every state and the mandatory pairs; the gate items of the rubric | approved contact sheet |
|
|
200
|
+
| `internal` | an admin panel, an internal tool, a CLI | none owed | recommended; the project linter is the floor | the functional look |
|
|
201
|
+
| `ad` | a creative that runs as an advertisement or a store asset | Brief, Mode, References, Markers, ADA (its rubric profile and safe zones), Open | **gate**: the ad profile and the safe zones | approved contact sheet |
|
|
202
|
+
|
|
203
|
+
Ask it as one question with a recommended answer read off the request — a landing or a
|
|
204
|
+
paywall is `flagship` unless the operator says otherwise — and never leave it to stage
|
|
205
|
+
5: a class decided at build time is decided by whoever wants the build to pass. A
|
|
206
|
+
work-graph node that builds the surface copies the class as `surface_class`, and
|
|
207
|
+
`graph.py certify` then owes the fourth, `visual` reading on `flagship` and `product`
|
|
208
|
+
([`certification.md`](certification.md) → *The fourth reading*).
|
|
191
209
|
- **Artifact:** lock the resolved decisions into a **task brief** committed at
|
|
192
|
-
`<artifacts>/specs/YYYY-MM-DD-<topic>-brief.md` (scope, users/UI verdict,
|
|
193
|
-
constraints, assumptions, explicitly-deferred items, done-criteria) **plus the
|
|
210
|
+
`<artifacts>/specs/YYYY-MM-DD-<topic>-brief.md` (scope, users/UI verdict and,
|
|
211
|
+
for a user-facing task, `surface_class`, constraints, assumptions, explicitly-deferred items, done-criteria) **plus the
|
|
194
212
|
autonomy sweep's per-stage answers and the model decision**. Seed it from
|
|
195
213
|
the skill's `templates/brief.md` skeleton — but only when absent, never
|
|
196
214
|
overwrite an existing brief. Stages 2–4 build on this brief; stages 5–10 read
|
|
@@ -214,7 +232,8 @@ never that the work was skipped quietly.
|
|
|
214
232
|
a recorded answer or an explicit deferral, **every answer that contradicted a
|
|
215
233
|
harvested source has a recorded resolution** (which governs, and whether the doc
|
|
216
234
|
is now stale), no open contradictions, **every
|
|
217
|
-
autonomy-sweep row is answered or explicitly marked "stop and ask here"**,
|
|
235
|
+
autonomy-sweep row is answered or explicitly marked "stop and ask here"**, **a
|
|
236
|
+
user-facing task's brief names its `surface_class`**, the
|
|
218
237
|
**REQ table is written and every row names its check**, the carry-over ledger is
|
|
219
238
|
seeded, **`.task-pipeline/run.md` exists and the header block has been printed**
|
|
220
239
|
([`progress.md`](progress.md)), the model decision is recorded, and the operator
|
|
@@ -314,8 +333,9 @@ never that the work was skipped quietly.
|
|
|
314
333
|
ssheleg/super-ux`). super-ux builds a traced chain — walk it top-down (see its
|
|
315
334
|
`system-map.md`):
|
|
316
335
|
0. **Destination first, when Figma is on.** The brief already names the team/org
|
|
317
|
-
and the
|
|
318
|
-
|
|
336
|
+
and the files — **one per surface** (App, Web, ASO: store screenshots, icon and
|
|
337
|
+
logo), never one file for every frame; `docs/ux/foundation.md` → *Design tooling*
|
|
338
|
+
is the canonical record. Confirm each **resolves** before drawing. **Never create a file while a
|
|
319
339
|
recorded one resolves; if it doesn't resolve, stop and ask — never create a
|
|
320
340
|
replacement** (that is the duplicate, and it hides a permissions problem).
|
|
321
341
|
A creation happens at most once, in the named team, and its URL is written to
|
|
@@ -350,7 +370,10 @@ never that the work was skipped quietly.
|
|
|
350
370
|
disagree together on one screen. The full doctrine — each track's scope and
|
|
351
371
|
out-of-scope, the refusal sentences, the four contradictions the check
|
|
352
372
|
catches — is [`spec.md`](spec.md) → *The COPY and VISUAL tracks, and their
|
|
353
|
-
convergence*, its one home.
|
|
373
|
+
convergence*, its one home. **The VISUAL track leaves a trace, not a fact**: a
|
|
374
|
+
director record at `docs/design/<surface>/director-record.md`, whose fields the gate
|
|
375
|
+
reads by the brief's `surface_class` (`python3 scripts/visual_gate.py record <file>
|
|
376
|
+
--class <c>`), and a refusal is that same file saying `Mode: declined` and why.
|
|
354
377
|
- **Spec:** write the approved design to
|
|
355
378
|
`<artifacts>/specs/YYYY-MM-DD-<topic>-design.md` and commit it. Lock all
|
|
356
379
|
shared contracts (types, schemas, signatures, file layout). For UI tasks the
|
|
@@ -367,12 +390,23 @@ never that the work was skipped quietly.
|
|
|
367
390
|
designed, validated and approved; scenarios validated in `docs/ux/scenarios.md`;
|
|
368
391
|
the linter passes; every user-facing spec requirement traces to a scenario ID
|
|
369
392
|
(or an explicit v1-mode/tiny-project waiver by the operator). **With Figma on:
|
|
370
|
-
the canonical record names one file
|
|
371
|
-
|
|
372
|
-
|
|
393
|
+
the canonical record names one file per surface (App, Web, ASO), and every
|
|
394
|
+
`screens.md` frame link's `:fileKey` is one of them** — a string match, not a
|
|
395
|
+
judgement (`python3 scripts/visual_gate.py filekeys --record docs/ux/foundation.md
|
|
396
|
+
--screens docs/ux/screens.md`); a key outside the set means the run drew in a file
|
|
397
|
+
nobody recorded and nobody will open. **Every user-facing string went
|
|
373
398
|
through the COPY track or the refusal is recorded**, and **the visual layer went
|
|
374
399
|
through the VISUAL track or the refusal is recorded** — a recorded refusal passes
|
|
375
400
|
this gate and an unmentioned one does not, which is the only difference that matters.
|
|
401
|
+
**And the VISUAL track is checked by its trace, not by the fact that it ran:** on a
|
|
402
|
+
`flagship`, `product` or `ad` surface the director record exists and carries the
|
|
403
|
+
fields its class owes — `python3 scripts/visual_gate.py record
|
|
404
|
+
docs/design/<surface>/director-record.md --class <surface_class>` exits 0. That
|
|
405
|
+
command checks the headings itself and runs the record's own validator from `sheleg-design`
|
|
406
|
+
(`--check-record`) where it is installed and new enough; where it is not, the
|
|
407
|
+
validator reads **NOT_RUN** beside the verdict — never PASS — and the gate stands on
|
|
408
|
+
the floor, said so. A record saying `Mode: declined` with its reason is the refusal,
|
|
409
|
+
and it passes.
|
|
376
410
|
**Where both tracks ran, their convergence check is recorded** — findings with the
|
|
377
411
|
ruling, or `Tracks converge: clean`; a screen where each track is right alone and they
|
|
378
412
|
disagree together is the defect neither track's own review can see.
|
|
@@ -435,7 +469,12 @@ never that the work was skipped quietly.
|
|
|
435
469
|
requires — never parked silently** — a browser finding filed without a ruling is the
|
|
436
470
|
diff-review verdict wearing a screenshot; the look was worth taking only if it can
|
|
437
471
|
still change the code or is on record as deliberately not doing so. Absent, say the surface was verified by reading
|
|
438
|
-
the diff and treat it as the weaker claim it is.
|
|
472
|
+
the diff and treat it as the weaker claim it is. **On a surface whose brief names a
|
|
473
|
+
`surface_class`, the project linter runs here too**, after each task that changes a
|
|
474
|
+
rendered surface — `python3 scripts/visual_gate.py lint <dir>`; an S1 finding is
|
|
475
|
+
fixed in the task, and NOT_RUN (exit 3) is recorded as such
|
|
476
|
+
([`browser.md`](browser.md) → *The visual half*). A slop marker caught while the
|
|
477
|
+
implementer is dispatched costs a line; caught on the contact sheet it costs a round. Stage 6 repeats this over the whole tree; this one catches it while the
|
|
439
478
|
implementer that wrote it is still dispatched. The matrix pointed this companion at
|
|
440
479
|
stages 5–6 from the day it was added and **this stage had never named it** — found by
|
|
441
480
|
the guard comparing the two, not by a reader.
|
|
@@ -472,7 +511,10 @@ never that the work was skipped quietly.
|
|
|
472
511
|
printed beside their floors; the **full** suite is green (not just the new tests); new/changed code
|
|
473
512
|
is covered; **every check this run added or widened has been probed both ways —
|
|
474
513
|
seen rejecting a planted defect and passing the clean tree, asserted on its exit
|
|
475
|
-
code** ([`probing.md`](probing.md)); no `skip`/`xfail` smuggling a red suite past the gate
|
|
514
|
+
code** ([`probing.md`](probing.md)); no `skip`/`xfail` smuggling a red suite past the gate; **on a
|
|
515
|
+
`flagship`, `product` or `ad` surface where the VISUAL track ran, the visual half's
|
|
516
|
+
`visual_gate.py sheet` exits 0** — NOT_RUN stops and asks, it is not green
|
|
517
|
+
([`browser.md`](browser.md) → *The visual half*). Never advance
|
|
476
518
|
to deploy on a red or partial run. **The carry-over count is printed beside this
|
|
477
519
|
verdict** — a ratchet nobody prints is a TODO with a better name
|
|
478
520
|
([`audit.md`](audit.md)) — **and so are the disclosures**, `abstained` and
|
|
@@ -502,6 +544,19 @@ never that the work was skipped quietly.
|
|
|
502
544
|
run that answers *the surface was checked* by pointing at its spec suite has answered
|
|
503
545
|
a different question. Where the suite is the thing that changed, the look is what
|
|
504
546
|
proves it runs against a page that renders.
|
|
547
|
+
- **The visual half is a separate check, and on most user-facing classes a gate.** The
|
|
548
|
+
look above reads the accessibility tree; it cannot say whether the surface looks like
|
|
549
|
+
what was designed. Where the stage-3 VISUAL track ran on a `flagship`, `product` or
|
|
550
|
+
`ad` surface, stage 6 also owes **the contact sheet** — frames over the `SCR-NN` states
|
|
551
|
+
× viewport × theme × text × locale (pairwise, plus the mandatory pairs), each with its
|
|
552
|
+
capture record, diffed against its Figma frame or approved baseline, the project
|
|
553
|
+
linter run, and the rubric read by a judge that is not the builder — and its command
|
|
554
|
+
exits 0: `python3 scripts/visual_gate.py sheet <contact-sheet.json> --class
|
|
555
|
+
<surface_class> --artifact-root <frames>` (NOT_RUN, exit 3, is not green). On
|
|
556
|
+
`internal` it is recommended and the linter is the floor. The re-render budget is
|
|
557
|
+
**one, two at most**, then `unresolved` to the person — never round three — and only
|
|
558
|
+
external, specific feedback starts a round. The procedure and the order of the checks:
|
|
559
|
+
[`browser.md`](browser.md) → *The visual half*.
|
|
505
560
|
- **What the look finds is fixed here.** A rendering defect found at stage 6 is a
|
|
506
561
|
stage-6 finding: fix it, look again, then call the stage green. Filing it to the
|
|
507
562
|
board and advancing is how a run reports *checked in a browser* for a page it has
|
|
@@ -672,7 +727,14 @@ never that the work was skipped quietly.
|
|
|
672
727
|
surface**: read it against the spec section that covers its `SCR-` id and against
|
|
673
728
|
what shipped. The super-ux linter proves a frame link exists, is named right and
|
|
674
729
|
is not stale — it cannot read the picture, so a frame promising a limit, a meter
|
|
675
|
-
or a tier nobody built passes every lint there is.
|
|
730
|
+
or a tier nobody built passes every lint there is. **Where the VISUAL track ran, the
|
|
731
|
+
walk carries one more row: visual intent ↔ final render** — the director record (the
|
|
732
|
+
brief's falsifier, the signature moment, the rubric written before any render) read
|
|
733
|
+
against the **approved contact sheet**. On a `flagship`, `product` or `ad` surface the
|
|
734
|
+
sheet is approved (`visual_gate.py sheet … --require-approval` exits 0), and the last
|
|
735
|
+
`review:` line per surface — how many human rounds it took — is copied into the
|
|
736
|
+
acceptance file, because `.task-pipeline/run.md` does not outlive the run
|
|
737
|
+
([`audit.md`](audit.md) → the `V→R` seam). An absence
|
|
676
738
|
becomes a **new REQ row with its check** and *then* the table is written;
|
|
677
739
|
appending after the table is how acceptance goes green over a gap. Findings that
|
|
678
740
|
belong to a lower layer go back to that layer (spec → stage 3, plan → stage 4).
|
|
@@ -59,7 +59,7 @@ conditional on the code, never merely sequenced after it.**
|
|
|
59
59
|
| `goal` | the release goal | `0` · `3` unstated |
|
|
60
60
|
| `add` | the id it allocated | `0` · `1` refused |
|
|
61
61
|
| `park` | the id and the reason | `0` · `1` refused |
|
|
62
|
-
| `certify` | the round, and on a failure every `breaks` finding with its fix and its check | `0`
|
|
62
|
+
| `certify` | the round, and on a failure every `breaks` finding with its fix and its check | `0` every owed tier passed — three, or four with `visual` on a flagship or product node · `1` a tier failed, a required one is missing, or a report is malformed |
|
|
63
63
|
| `close` | the goal, the new frontier count, and what was not verified | `0` · `1` refused **or the verdict stops the run** |
|
|
64
64
|
| `producer` | what produced this proof — actor, model, runtime, skill, config, commit, trace | `0` |
|
|
65
65
|
| `doctrine` | how many of the bundle's reference files this run opened | `0` |
|
|
@@ -84,7 +84,7 @@ A **`parked`** node is the single exemption: it is the one node nobody will clos
|
|
|
84
84
|
*n/a — parked* in that field is confidence without correctness. `park` never removes what
|
|
85
85
|
the node said it would run.
|
|
86
86
|
|
|
87
|
-
**A node is closed by three readings, not one.** `certify` takes one tier report from each of `unit`, `seam` and `product` — dispatched blind and in parallel — requires all three to pass, and assembles the seven-key verdict `close` consumes. `close`'s contract is unchanged; what changed is that the verdict is now built from three readings at different distances instead of written from one, because a change can be correct where it was made and wrong one level out. A failing round records itself and leaves the node open. Doctrine: [`certification.md`](certification.md).
|
|
87
|
+
**A node is closed by three readings, not one.** `certify` takes one tier report from each of `unit`, `seam` and `product` — dispatched blind and in parallel — requires all three to pass, and assembles the seven-key verdict `close` consumes. `close`'s contract is unchanged; what changed is that the verdict is now built from three readings at different distances instead of written from one, because a change can be correct where it was made and wrong one level out. A failing round records itself and leaves the node open. **A node whose `surface_class` is `flagship` or `product` owes a fourth, `visual` report** — the contact sheet read against the director record — and `certify` refuses the round without it; on any other node a `visual` report is accepted when given. Doctrine: [`certification.md`](certification.md).
|
|
88
88
|
|
|
89
89
|
**`close` stamps the commit; the verifier never supplies it.** A verdict written after the
|
|
90
90
|
tree moved is evidence about a different tree, and an agent cannot name the wrong commit if
|