task-pipeline-skill 1.88.1 → 1.90.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +109 -0
- package/README.md +2 -2
- package/SKILL-CARD.md +1 -1
- package/package.json +3 -3
- package/plugins/task-pipeline/.claude-plugin/plugin.json +1 -1
- package/plugins/task-pipeline/agents/verifier-product.md +3 -1
- package/plugins/task-pipeline/agents/verifier-visual.md +115 -0
- package/plugins/task-pipeline/agents/verifier.md +2 -1
- package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +5 -5
- package/plugins/task-pipeline/skills/task-pipeline/graph.schema.json +22 -1
- package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +4 -4
- package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +12 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/audit.md +14 -8
- package/plugins/task-pipeline/skills/task-pipeline/references/browser.md +107 -3
- package/plugins/task-pipeline/skills/task-pipeline/references/build.md +41 -2
- package/plugins/task-pipeline/skills/task-pipeline/references/certification.md +46 -5
- package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +29 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/conventions.md +3 -3
- package/plugins/task-pipeline/skills/task-pipeline/references/doctrine-map.md +1 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/grill.md +15 -9
- package/plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md +23 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/portability.md +1 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/spec.md +27 -3
- package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +83 -13
- package/plugins/task-pipeline/skills/task-pipeline/references/work-graph.md +2 -2
- package/plugins/task-pipeline/skills/task-pipeline/scripts/graph.py +45 -16
- package/plugins/task-pipeline/skills/task-pipeline/scripts/stage_checkpoint.py +21 -1
- package/plugins/task-pipeline/skills/task-pipeline/scripts/visual_gate.py +728 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/brief.md +5 -2
- package/plugins/task-pipeline/skills/task-pipeline/templates/browser-claims.json +223 -1
- package/plugins/task-pipeline/skills/task-pipeline/templates/run.md +10 -0
|
@@ -28,6 +28,7 @@ which is how a run says *I checked the browser* and means *I ran the unit tests*
|
|
|
28
28
|
- Sessions, and why an agent needs them
|
|
29
29
|
- Reading a look vs gating on one
|
|
30
30
|
- "Tested in a browser" is three different claims
|
|
31
|
+
- The visual half — pixels against intent, under a contract
|
|
31
32
|
- Getting past a login, and past a backend
|
|
32
33
|
- When the look finds something: debugging the spec that missed it
|
|
33
34
|
- Evidence a reader can open
|
|
@@ -48,7 +49,9 @@ Three consequences the doctrine rests on:
|
|
|
48
49
|
|
|
49
50
|
- **A look costs a page of text and no vision model.** This is why the pipeline can ask
|
|
50
51
|
for one at three stages without the cost being an argument. A `screenshot` exists in
|
|
51
|
-
both channels and *is* pixels — take one for a human
|
|
52
|
+
both channels and *is* pixels — in the functional look, take one for a human, not for
|
|
53
|
+
you to read. Reading pixels against the design's intent is a different check with its
|
|
54
|
+
own contract: *The visual half*, below.
|
|
52
55
|
- **The ref is a fact about the page as rendered**, so `click e12` after a snapshot is
|
|
53
56
|
deterministic in a way a coordinate never is.
|
|
54
57
|
- **A ref that no longer resolves is a finding, not an error to retry past.** The element
|
|
@@ -125,7 +128,7 @@ commonest way a run reports a green it does not have.
|
|
|
125
128
|
|
|
126
129
|
| | What it is | What it proves | Where it counts |
|
|
127
130
|
|---|---|---|---|
|
|
128
|
-
| **The look** | an agent driving a page: open, snapshot, console, network | that this surface renders, right now, and what the browser said while it did | the **look**, stage 6 — recommended, never a gate |
|
|
131
|
+
| **The look** | an agent driving a page: open, snapshot, console, network | that this surface renders, right now, and what the browser said while it did | the **functional look**, stage 6 — recommended, never a gate. Its **visual half** is a gate on a `flagship`, `product` or `ad` surface — *The visual half*, below |
|
|
129
132
|
| **The spec suite** | `playwright test` — the **test runner** | that the assertions someone wrote still hold, on the paths someone thought to write | the **suite** half of the stage-6 gate, counted with every other test |
|
|
130
133
|
| **The library** | `require('playwright')` — `chromium`/`firefox`/`webkit`, `devices`, `request`, `selectors` | whatever your own script asserts; it is an automation API, not a test framework | wherever the project already runs it |
|
|
131
134
|
|
|
@@ -155,6 +158,105 @@ in the claim's own state (the initial screenshot closes nothing about opened/err
|
|
|
155
158
|
a suite PASS closes no look claim; a toggle owes its full cycle; no browser channel
|
|
156
159
|
is NOT_RUN with the reason.
|
|
157
160
|
|
|
161
|
+
## The visual half — pixels against intent, under a contract
|
|
162
|
+
|
|
163
|
+
Everything above reads the **accessibility tree**: it proves the surface renders and the
|
|
164
|
+
browser stayed quiet. It cannot say whether the surface looks like what was designed —
|
|
165
|
+
the wrong weight on a heading, a card that lost its spacing at 200 % text, a dark theme
|
|
166
|
+
that drops a border, a frame that drifted from the Figma it was built from. That is a
|
|
167
|
+
different question, and asking it of a snapshot is how a run reports *it looks right*
|
|
168
|
+
having read no pixels at all. The **functional look** and the **visual half** are two
|
|
169
|
+
checks; neither discharges the other.
|
|
170
|
+
|
|
171
|
+
**When it runs, and when it gates — by the brief's `surface_class`**
|
|
172
|
+
([`stages.md`](stages.md) → stage 0, *The surface class*):
|
|
173
|
+
|
|
174
|
+
| Class | Visual half at stages 5–6 | Human pass at stage 10 |
|
|
175
|
+
|---|---|---|
|
|
176
|
+
| `flagship` | **gate** — full matrix, pairwise across every axis, judge items per the full rubric | **gate** — the approved contact sheet |
|
|
177
|
+
| `product` | **gate** — every state, the mandatory pairs; pairwise holes reported | **gate** — the approved contact sheet |
|
|
178
|
+
| `ad` | **gate** — the ad rubric profile and the safe zones | **gate** — the approved contact sheet |
|
|
179
|
+
| `internal` | recommended — the project linter is the floor; a sheet, if made, is checked for honesty | the functional look closes it |
|
|
180
|
+
|
|
181
|
+
It gates only where the stage-3 VISUAL track ran. A recorded refusal (*«без дизайна»*,
|
|
182
|
+
`Mode: declined`) turns the visual half into the functional look plus the linter, said in
|
|
183
|
+
the close-out — the same rule as every other refusal in the pipeline.
|
|
184
|
+
|
|
185
|
+
**The checks run cheapest first, and the person last:**
|
|
186
|
+
|
|
187
|
+
1. **Deterministic.** The project linter — `python3 scripts/visual_gate.py lint <dir>`,
|
|
188
|
+
which runs sheleg-design's `--lint` where it is installed and answers **NOT_RUN (exit 3)**
|
|
189
|
+
where it is not — plus axe or Lighthouse, the token check, the type checker. An S1
|
|
190
|
+
finding blocks, whatever anyone says about the picture later. **With Figma on, the
|
|
191
|
+
token check includes the drift probe**: `python3 scripts/visual_gate.py tokens
|
|
192
|
+
--figma <variables.json> --css <tokens.css>` compares the variable names exported
|
|
193
|
+
from the file (the JSON `get_variable_defs` returns, or the REST export) with the
|
|
194
|
+
custom properties the pack's token file declares. It names each variable with no
|
|
195
|
+
property, each property with no variable, and each variable whose WEB code syntax
|
|
196
|
+
is not the property the file actually declares. A name maps by kebab-case
|
|
197
|
+
(`Color/Text Muted` → `--color-text-muted`) unless code syntax says otherwise.
|
|
198
|
+
Exit `0` PASS · `1` FAIL · `2` unreadable · `3` NOT_RUN. No export is NOT_RUN,
|
|
199
|
+
never PASS: save the export beside the frames and run again, or record the probe
|
|
200
|
+
as not run.
|
|
201
|
+
2. **Regression** against the approved baseline, where one exists (`toHaveScreenshot`, or
|
|
202
|
+
the platform's snapshot test).
|
|
203
|
+
3. **The matrix.** One frame per `SCR-NN` state the screen map lists (default, loading,
|
|
204
|
+
empty, error, offline, long-content, keyboard-up, first-run — the ones that apply),
|
|
205
|
+
across viewport × theme × text size × locale: **pairwise coverage** — every value of one
|
|
206
|
+
axis meets every value of every other in some frame — plus the **mandatory pairs**
|
|
207
|
+
(dark × large text, RTL × narrow). Never the full cross product: every state at every
|
|
208
|
+
combination of every axis value is hundreds of frames, and nobody reads them. Seed the
|
|
209
|
+
worst-case data first (`break-ui`, [`companion-skills.md`](companion-skills.md) →
|
|
210
|
+
*Visual lanes*) so the frames show long names and empty lists, not the demo account.
|
|
211
|
+
4. **Each frame carries its capture record** — revision, route, state, viewport, locale,
|
|
212
|
+
theme, motion, captured-at, source — and is disqualified, not passed, when it is blank,
|
|
213
|
+
stale, of the wrong route or state, or taken before the fonts rendered. This is
|
|
214
|
+
sheleg-design's visual-review contract; the pipeline only refuses a frame without it.
|
|
215
|
+
Where a Figma frame or an approved baseline exists, the frame is **diffed against it**,
|
|
216
|
+
and a failing diff is never a PASS the run writes: either the build drifted, or the
|
|
217
|
+
person approves a new baseline.
|
|
218
|
+
5. **The rubric**, read by a judge that is not the agent that built the surface: a
|
|
219
|
+
checklist per task, never a single score; pairwise only against the approved
|
|
220
|
+
reference and in both orders; three samples, and disagreement is `uncertain`, which
|
|
221
|
+
goes to the person rather than to a coin. **A judge item (J) is `NOT_ASSESSED` until a
|
|
222
|
+
labelled set exists and the judge's agreement with it is measured** — a verdict from an
|
|
223
|
+
uncalibrated judge is a guess with a format. A gate item (G) is deterministic and the
|
|
224
|
+
judge never overrides its FAIL. Every FAIL is a triple: *region → defect → fix*.
|
|
225
|
+
6. **One human pass** over the contact sheet. Approval makes its frames the next baseline.
|
|
226
|
+
|
|
227
|
+
**The contact sheet is the one surface the person reviews**, and its data is
|
|
228
|
+
`templates/browser-claims.json` — the same file, **not a second schema**. A look row
|
|
229
|
+
that carries `axes` is a frame: `axes{viewport,theme,text,locale}`, `capture{revision,
|
|
230
|
+
route,motion,captured_at,source}`, `figma_frame`, `baseline`, `diff`, `rubric[]`. The file
|
|
231
|
+
names its `surface`, `revision`, `review_rounds`, `approved_by` and `approved_at`. The
|
|
232
|
+
viewing page is a self-contained local HTML beside the frames — a full document, `<meta
|
|
233
|
+
charset="utf-8">`, no CDN. **Frames and the HTML stay out of git; the JSON goes in.**
|
|
234
|
+
|
|
235
|
+
```bash
|
|
236
|
+
python3 scripts/visual_gate.py sheet design/review/contact-sheet.json \
|
|
237
|
+
--class product --artifact-root design/review --states SCR-01/default,SCR-01/empty
|
|
238
|
+
# stage 10 adds --require-approval
|
|
239
|
+
```
|
|
240
|
+
|
|
241
|
+
Exit `0` PASS · `1` FAIL · `2` unreadable · `3` NOT_RUN — a frame that did not run makes
|
|
242
|
+
the whole verdict NOT_RUN, never PASS. **On a gated class NOT_RUN stops the stage and
|
|
243
|
+
asks**: connect a capture channel, or the operator accepts the surface as `unverified`,
|
|
244
|
+
recorded in the brief and named in the close-out. It is never reported around, which is
|
|
245
|
+
the failure `DEC-0004` kept the functional look ungated to avoid; `DEC-0006` gates the
|
|
246
|
+
visual half on these classes and keeps that exit open in words.
|
|
247
|
+
|
|
248
|
+
**Two rules keep the loop from becoming the work** ([`loop-guard.md`](loop-guard.md) →
|
|
249
|
+
*The re-render loop*): a re-render budget of **one, two at most**, after which a failing
|
|
250
|
+
item goes to the person as `unresolved`; and **only external, specific feedback** starts a
|
|
251
|
+
round — a linter line, a diff, an audit item, a triple — never "look again". Each return
|
|
252
|
+
from the person adds one to `review_rounds` and one `review:` line to the run ledger, so the
|
|
253
|
+
number of passes a surface took is measured, not remembered.
|
|
254
|
+
|
|
255
|
+
**Native surfaces are not web surfaces.** A web render styled as a phone is a mockup. A
|
|
256
|
+
native screen's frames come from a simulator or a device (XCUITest snapshots, Compose
|
|
257
|
+
screenshot tests) and its accessibility from the platform's own audit; with neither, the
|
|
258
|
+
native rows stay `NOT_RUN`, said in words.
|
|
259
|
+
|
|
158
260
|
## Getting past a login, and past a backend
|
|
159
261
|
|
|
160
262
|
A surface behind auth is the usual reason a run skips the look. Both channels solve it,
|
|
@@ -277,7 +379,9 @@ On a CI box, headed is the failure you will spend an hour on.
|
|
|
277
379
|
| The excuse | Why it fails |
|
|
278
380
|
|---|---|
|
|
279
381
|
| *"`playwright test` is green, the surface is checked."* | The suite asserts what someone wrote down. `DEC-0004`: it is the coverage half, never the look. |
|
|
280
|
-
| *"I took a screenshot, so I looked."* | A screenshot is pixels you did not read. The look is `snapshot` + `console` + `requests`, and the verdict quotes them. |
|
|
382
|
+
| *"I took a screenshot, so I looked."* | A screenshot is pixels you did not read. The functional look is `snapshot` + `console` + `requests`, and the verdict quotes them; reading the pixels is the visual half, with a capture record per frame and a matrix behind it. |
|
|
383
|
+
| *"The snapshot is clean, so it looks right."* | The tree says the heading exists, not that it rendered at the right weight, in the dark theme, at 200 % text. That is the visual half's question, and on a flagship, product or ad surface it is a gate. |
|
|
384
|
+
| *"The judge said it looks great."* | An uncalibrated judge's J items are `NOT_ASSESSED`, and no judge overrides a deterministic FAIL. A verdict needs a checklist, a reference, both orders and three samples. |
|
|
281
385
|
| *"The click failed, I'll find a better selector."* | A ref that stopped resolving **is the finding**. Re-snapshot and report what moved. |
|
|
282
386
|
| *"The docs say the CLI has no `tracing`."* | A vendor page is a claim; `--help` is the tool. This file was written against `--help` **because** a page-derived claim shipped here and was wrong. |
|
|
283
387
|
| *"The tool list is in the docs."* | The page listed tools this version does not ship, and omitted that tracing, video and PDF need `--caps`. Ask the server: 24 tools default, 42 with all caps. |
|
|
@@ -532,10 +532,13 @@ it invents what was already decided.
|
|
|
532
532
|
approximately, and approximate is indistinguishable from exact in a report.
|
|
533
533
|
3. **A component with a Code Connect mapping is not rewritten.** If
|
|
534
534
|
`get_code_connect_map` names a code component for that node, the screen uses it.
|
|
535
|
-
Reimplementing it is a silent fork of the design system.
|
|
535
|
+
Reimplementing it is a silent fork of the design system. *Code Connect, kept*, below,
|
|
536
|
+
is what keeps that mapping true.
|
|
536
537
|
4. **A token names its variable.** A raw hex or px where the file has a variable is a
|
|
537
538
|
token that has quietly split in two. `get_variable_defs` is the canon; a screenshot
|
|
538
|
-
is a way to *look*, never a way to *know*.
|
|
539
|
+
is a way to *look*, never a way to *know*. Whether the names still agree is
|
|
540
|
+
measured, not assumed: `visual_gate.py tokens` ([`browser.md`](browser.md) →
|
|
541
|
+
*The visual half*).
|
|
539
542
|
5. **The frame is a contract at its own width.** It is one width and said nothing about
|
|
540
543
|
the others, so behaviour at other breakpoints — and states the frame does not draw,
|
|
541
544
|
like error, empty and loading — is a **decision that gets recorded**, not guessed.
|
|
@@ -559,6 +562,42 @@ same false confidence as an unproven green.
|
|
|
559
562
|
platform constraint, an accessibility floor, a breakpoint — write what and why.
|
|
560
563
|
Otherwise *"built from the frame"* and *"built to look like it"* read identically.
|
|
561
564
|
|
|
565
|
+
### Code Connect, kept
|
|
566
|
+
|
|
567
|
+
Rule 3 trusts the mapping. It is only worth trusting while it stays true. A mapping
|
|
568
|
+
that still names a component's old props hands the next agent a snippet that does
|
|
569
|
+
not compile, so it is followed by reimplementing the component. That is the fork
|
|
570
|
+
rule 3 exists to prevent, reached by obeying it.
|
|
571
|
+
|
|
572
|
+
- **A component whose API changes updates its Code Connect mapping in the same
|
|
573
|
+
change.** That covers a renamed or removed prop, a new variant, or a moved import
|
|
574
|
+
path. The mapping file is part of the component's diff and reviewed with it. A
|
|
575
|
+
mapping left for later is wrong from the next merge onwards. Figma's own guidance
|
|
576
|
+
says the same: *"When component APIs change in your codebase, update the
|
|
577
|
+
corresponding Code Connect mappings"* ([Code Connect integration][fcc], accessed
|
|
578
|
+
2026-10-08).
|
|
579
|
+
- **A core component with no mapping gets an offer, not a silent skip.** Where a
|
|
580
|
+
task builds on a design-system component that the file draws and Code Connect does
|
|
581
|
+
not map, the run offers to map it with `/figma-code-connect`, Figma's public
|
|
582
|
+
plugin skill ([`companion-skills.md`](companion-skills.md)). The offer names which
|
|
583
|
+
components and which file. Mapping publishes to the Figma file, so it runs only on
|
|
584
|
+
an explicit go, like drawing a missing frame. Absent the skill, the run names the
|
|
585
|
+
unmapped components in the close-out and goes on.
|
|
586
|
+
- **Never rewrite a mapped component instead of using it** — rule 3, unchanged. An
|
|
587
|
+
API that no longer fits the frame is a change to the component and its mapping
|
|
588
|
+
together, never a local copy.
|
|
589
|
+
|
|
590
|
+
Why it pays, in Figma's words and on Figma's measurement: their own evals report *"a
|
|
591
|
+
19.6% reduction in median task duration, a full point of improvement in code quality
|
|
592
|
+
on a 1–4 scale, and a 29.5% reduction in token usage"* with Code Connect
|
|
593
|
+
([Figma MCP use cases][fuc], accessed 2026-10-08). **That is Figma's unaudited
|
|
594
|
+
number** — a vendor's eval of its own feature, with no published protocol here and
|
|
595
|
+
not reproduced by this pipeline. Quote it as theirs, never as a measured property
|
|
596
|
+
of a run.
|
|
597
|
+
|
|
598
|
+
[fcc]: https://developers.figma.com/docs/figma-mcp-server/code-connect-integration/
|
|
599
|
+
[fuc]: https://www.figma.com/resource-library/figma-mcp-use-cases/
|
|
600
|
+
|
|
562
601
|
## 5. Final whole-branch review
|
|
563
602
|
|
|
564
603
|
After the last task: build a package over `MERGE_BASE`..`HEAD`
|
|
@@ -15,6 +15,7 @@ queue is — graph or plan — is what decides, not the mood of the closer.
|
|
|
15
15
|
|
|
16
16
|
- Why one verifier is not enough, stated as the failure it produces
|
|
17
17
|
- The three tiers
|
|
18
|
+
- The fourth reading — `visual`, on a flagship or product surface
|
|
18
19
|
- Blind, and it is the whole design
|
|
19
20
|
- A pass has to mean something, so two rules have teeth
|
|
20
21
|
- The report, and where each field lands
|
|
@@ -60,7 +61,42 @@ that reads no code is not the soft one.
|
|
|
60
61
|
|
|
61
62
|
Agents: [`../../../agents/verifier-unit.md`](../../../agents/verifier-unit.md),
|
|
62
63
|
[`verifier-seam.md`](../../../agents/verifier-seam.md),
|
|
63
|
-
[`verifier-product.md`](../../../agents/verifier-product.md)
|
|
64
|
+
[`verifier-product.md`](../../../agents/verifier-product.md) — and, on a visual surface,
|
|
65
|
+
[`verifier-visual.md`](../../../agents/verifier-visual.md), below.
|
|
66
|
+
|
|
67
|
+
## The fourth reading — `visual`, on a flagship or product surface
|
|
68
|
+
|
|
69
|
+
The three tiers read code, what reaches it, and what the product says about it. **None
|
|
70
|
+
of them opens a picture**, and on a surface whose look is part of the requirement that
|
|
71
|
+
leaves a whole level unread: a node can pass unit, seam and product while its empty
|
|
72
|
+
state renders grey on grey at 200 % text, and every report is truthful about what its
|
|
73
|
+
tier saw.
|
|
74
|
+
|
|
75
|
+
| Tier | Subject | Characteristic finding |
|
|
76
|
+
|---|---|---|
|
|
77
|
+
| `visual` | the contact sheet, the director record, the project linter's output, the rubric | a frame that contradicts the record's intent; a gate item failing under a PASS; a hole in the state × axes matrix; a judge verdict from an uncalibrated judge |
|
|
78
|
+
|
|
79
|
+
**When it runs.** A node that builds a user-facing surface copies the brief's
|
|
80
|
+
`surface_class` ([`stages.md`](stages.md) → stage 0, *The surface class*). On
|
|
81
|
+
**`flagship` and `product`** the `visual` report is **required** — `graph.py certify`
|
|
82
|
+
refuses the round without it and names the tier. On `internal` and `ad` it is accepted
|
|
83
|
+
when given and counted like any other tier, and not demanded: an internal tool's floor
|
|
84
|
+
is the linter, and an ad's gate is its rubric profile on the contact sheet.
|
|
85
|
+
|
|
86
|
+
**It is the fourth blind reading, not a reviewer with a vision model.** Same eight-key
|
|
87
|
+
report, same `breaks`/`risk`, same blindness — it never sees the other three reports
|
|
88
|
+
and they never see it. Its own rules, because a judge of pixels fails in its own ways:
|
|
89
|
+
|
|
90
|
+
- **A checklist per task, never a score.** It answers the rubric's binary items against
|
|
91
|
+
the record and the sheet; "looks polished" is not an item.
|
|
92
|
+
- **Pairwise only against the approved reference, and in both orders**; three samples.
|
|
93
|
+
A verdict that flips with the order, or between samples, is **`uncertain`** and goes to
|
|
94
|
+
the person — it is neither a pass nor a `breaks`.
|
|
95
|
+
- **It never overrides a deterministic FAIL.** A gate item (G) or a linter S1 that
|
|
96
|
+
failed is a `breaks` whatever the picture looks like to it.
|
|
97
|
+
- **A judge item (J) is `NOT_ASSESSED` until a labelled set exists** and the judge's
|
|
98
|
+
agreement with it has been measured. It goes in `not_examined`, which reaches the
|
|
99
|
+
closing verdict as `not_verified` — the honest name for an opinion nobody calibrated.
|
|
64
100
|
|
|
65
101
|
## Blind, and it is the whole design
|
|
66
102
|
|
|
@@ -71,8 +107,10 @@ will paraphrase it back as product truth. The disagreement between blind reading
|
|
|
71
107
|
the instrument, so `graph.py certify` refuses a report whose prose cites another
|
|
72
108
|
tier's verdict.
|
|
73
109
|
|
|
74
|
-
Dispatch all three in one message so they run concurrently
|
|
75
|
-
its `serves`, and the diff —
|
|
110
|
+
Dispatch all three in one message so they run concurrently — all four on a visual
|
|
111
|
+
node. Give each the node id, its `serves`, and the diff — the `visual` tier also the
|
|
112
|
+
paths of the contact sheet, the director record and the linter output — nothing else,
|
|
113
|
+
and never another tier's output.
|
|
76
114
|
|
|
77
115
|
**The second axis, and it is the one an optimisation removes first: whoever produced
|
|
78
116
|
the fix never grades it.** Tier blindness is horizontal — no tier reads another's
|
|
@@ -146,6 +184,9 @@ consumer refuses.
|
|
|
146
184
|
# three reports in, one verdict out — exits 1 if any tier failed
|
|
147
185
|
graph.py certify --node N-007 \
|
|
148
186
|
--tier unit.json --tier seam.json --tier product.json
|
|
187
|
+
# a node with surface_class flagship or product: four, or the round is refused
|
|
188
|
+
graph.py certify --node N-008 \
|
|
189
|
+
--tier unit.json --tier seam.json --tier product.json --tier visual.json
|
|
149
190
|
|
|
150
191
|
# unchanged, and still the only thing that moves the graph
|
|
151
192
|
graph.py close --verdict .task-pipeline/verdict-N-007.json
|
|
@@ -176,8 +217,8 @@ has failed **every** round. A run spinning on one level needs the operator to se
|
|
|
176
217
|
|
|
177
218
|
## What this costs, said out loud
|
|
178
219
|
|
|
179
|
-
Three agents per node instead of one
|
|
180
|
-
paid per node rather than per run. The three are dispatched in parallel, so the
|
|
220
|
+
Three agents per node instead of one — four on a flagship or product surface. That is
|
|
221
|
+
the price of the visibility, and it is paid per node rather than per run. The three are dispatched in parallel, so the
|
|
181
222
|
wall-clock cost is roughly one reading; the token cost is three. A node whose
|
|
182
223
|
`check` is mechanical and whose blast radius is genuinely nil still pays it — and a
|
|
183
224
|
tier with nothing to find says so in `scope` and `not_examined` rather than being
|
|
@@ -24,11 +24,15 @@ one must never look alike.
|
|
|
24
24
|
> stop at the first that answers. The step stays **recommended and never a gate**: a gate
|
|
25
25
|
> an environment cannot satisfy is one an agent learns to report around, and *verified by
|
|
26
26
|
> reading the diff* already prices the absence honestly (`docs/DECISIONS.md`).
|
|
27
|
+
> **`DEC-0006`** scopes that to the functional look: the **visual half** is a gate on a
|
|
28
|
+
> `flagship`, `product` or `ad` surface, and its absent tool reads NOT_RUN and stops to
|
|
29
|
+
> ask rather than passing ([`browser.md`](browser.md) → *The visual half*).
|
|
27
30
|
|
|
28
31
|
## Contents
|
|
29
32
|
|
|
30
33
|
- Built in — nothing to install
|
|
31
34
|
- The matrix
|
|
35
|
+
- Visual lanes — tools, never entry points
|
|
32
36
|
- Optional bridge — substituting an external skill set
|
|
33
37
|
- Preflight (emit before stage 0)
|
|
34
38
|
- Is this skill itself current?
|
|
@@ -76,6 +80,31 @@ one must never look alike.
|
|
|
76
80
|
| ~~superpowers~~ | — | **Not a dependency.** Stages 2/4/5/6 run on the built-in doctrine above. See *Optional bridge* | — |
|
|
77
81
|
| ~~grill-me / grilling~~ | — | **Not a dependency.** The stage-0 grill is built in (`references/grill.md`) | — |
|
|
78
82
|
|
|
83
|
+
## Visual lanes — tools, never entry points
|
|
84
|
+
|
|
85
|
+
The visual half ([`browser.md`](browser.md) → *The visual half*) asks for checks no
|
|
86
|
+
companion above owns alone. These are the public tools that do each one. **Each is a
|
|
87
|
+
tool a stage reaches for, never an entry point and never a second route**: the stage
|
|
88
|
+
decides when, the tool does one job, and the close-out names which one it took. None is
|
|
89
|
+
required; an absent one is a check recorded as NOT_RUN with its reason.
|
|
90
|
+
|
|
91
|
+
| Tool | What it is for | Where in the stages |
|
|
92
|
+
|---|---|---|
|
|
93
|
+
| `break-ui` | seeds worst-case data — the longest name, the empty list, the 4-digit badge, the slow response — so frames show what a real account shows, not the demo | stage 6, **before** the contact-sheet screenshots |
|
|
94
|
+
| `review-animations`, `improve-animations` | reviews motion against its intent: timing, easing, interruption, reduced-motion fallback | stage 6, on a surface with motion; findings as triples |
|
|
95
|
+
| `mobile-native` | the mobile-web surface: viewport, safe areas, touch targets, keyboard-up states | stages 5–6, mobile-web frames of the matrix |
|
|
96
|
+
| `animate-expo` | React Native and Expo motion, built and reviewed on the platform's own primitives | stage 5, an RN or Expo surface |
|
|
97
|
+
| `webapp-testing`, `chrome-devtools` (`take_screenshot`, `lighthouse_audit`) | the frames of the matrix at each viewport and theme, and a Lighthouse pass as the deterministic floor | stage 6, the capture and the floor; stage 8 on a deployed target |
|
|
98
|
+
| `accessibility-review`, `a11y-debugging` | accessibility review of the rendered surface and debugging what it finds | stage 6, beside axe or Lighthouse |
|
|
99
|
+
| XCUITest `performAccessibilityAudit` | the iOS platform's own accessibility audit, run in UI tests | stage 6, a native iOS surface |
|
|
100
|
+
| Compose `enableAccessibilityChecks` | Android's accessibility checks in Compose UI tests | stage 6, a native Android surface |
|
|
101
|
+
| Playwright `toHaveScreenshot` | screenshot regression against the approved baseline, per state and theme | stage 6 (regression), and the nightly or release baseline after |
|
|
102
|
+
| axe-core | the deterministic accessibility floor of a web surface | stage 6, first in the cheap-first order |
|
|
103
|
+
| `figma-code-connect` (Figma's public plugin skill) | maps a design-system component in the file to its code component, so `get_design_context` hands the agent the real import instead of approximated markup | stage 5, **offered** for a core component the file draws and nothing maps — publishes to the file, so it needs a go ([`build.md`](build.md) → *Code Connect, kept*) |
|
|
104
|
+
|
|
105
|
+
A native screen is captured on a simulator or a device; a web render styled as a phone is
|
|
106
|
+
a mockup, and the native rows of the sheet stay NOT_RUN without one.
|
|
107
|
+
|
|
79
108
|
## Optional bridge — substituting an external skill set
|
|
80
109
|
|
|
81
110
|
An operator who already runs an equivalent skill set may map it onto stages 2/4/5/6
|
|
@@ -103,9 +103,9 @@ closing a stage with an unread CI verdict.
|
|
|
103
103
|
- Host self-update rules (module docs, runbooks, agent-self cards, etc.) — update
|
|
104
104
|
in the same change. Fix dangling links.
|
|
105
105
|
- **The design destination, on a project with no `docs/ux/`.** When the work uses
|
|
106
|
-
Figma but super-ux isn't in play, there is no `foundation.md` to hold the
|
|
107
|
-
the brief is canonical — and a brief is per-run. Write the team and
|
|
108
|
-
into the host's own docs (`CLAUDE.md`, or the README) in this change, so the next
|
|
106
|
+
Figma but super-ux isn't in play, there is no `foundation.md` to hold the files, so
|
|
107
|
+
the brief is canonical — and a brief is per-run. Write the team and each surface's
|
|
108
|
+
file URL into the host's own docs (`CLAUDE.md`, or the README) in this change, so the next
|
|
109
109
|
run reads the destination instead of creating a second file
|
|
110
110
|
([`grill.md`](grill.md) → *The design destination*).
|
|
111
111
|
- **The code graph:** [graphify](https://github.com/Graphify-Labs/graphify) —
|
|
@@ -30,7 +30,7 @@ UX track on a user-facing task.
|
|
|
30
30
|
| 3 Spec | `references/spec.md` |
|
|
31
31
|
| 4 Plan | `references/planning.md` |
|
|
32
32
|
| the queue the loop walks | `references/work-graph.md` |
|
|
33
|
-
| 5–8 · how a **work-graph node** is CLOSED — three blind readings at three distances, all three required (ceiling 3); a **prose-plan task** closes through `review.md` instead — one reviewer, five-round cap | `references/certification.md` |
|
|
33
|
+
| 5–8 · how a **work-graph node** is CLOSED — three blind readings at three distances, all three required, plus a fourth `visual` reading on a flagship or product surface (ceiling 3); a **prose-plan task** closes through `review.md` instead — one reviewer, five-round cap | `references/certification.md` |
|
|
34
34
|
| 5 Build (worktree, subagents, fix loop) | `references/build.md` + `references/review.md` |
|
|
35
35
|
| 5–6 TDD + suite gate | `references/tdd.md` |
|
|
36
36
|
| 5, 6, 8 The browser — the look, the spec suite, and the difference | `references/browser.md` |
|
|
@@ -18,7 +18,7 @@ coming back to the operator.
|
|
|
18
18
|
- Phase 2 — the gap check, then the loop
|
|
19
19
|
- Domain awareness
|
|
20
20
|
- The autonomy sweep
|
|
21
|
-
- The design destination — one file, decided here, never invented later
|
|
21
|
+
- The design destination — one file per surface, decided here, never invented later
|
|
22
22
|
- The REQ spine — the grill's other hard output
|
|
23
23
|
- Output
|
|
24
24
|
|
|
@@ -170,9 +170,9 @@ explicit "stop and ask me here":
|
|
|
170
170
|
| 0 Docs regime | where settled things live (the decision home — **one** per project, and an existing `docs/adr/` **is** it), who may write it, whether a lease mechanism is present or the run is `ungated`, the gate command and its ratchet floors, and whether this run may raise a floor ([`documentation.md`](documentation.md)) |
|
|
171
171
|
| 1 Docs | external libs/APIs/SDKs in play; any private ones context7 can't resolve → where their docs live |
|
|
172
172
|
| 2 Decompose | is this a platform (several capabilities/surfaces) or one module? if platform: deploy cadence — per module or once at the end |
|
|
173
|
-
| 2–3 Spec | UI verdict (arms super-ux); any scenario-tracing waiver |
|
|
173
|
+
| 2–3 Spec | UI verdict (arms super-ux) **and the surface class** — `flagship`, `product`, `internal` or `ad`, which selects the visual gate profile ([`stages.md`](stages.md) → stage 0, *The surface class*); any scenario-tracing waiver |
|
|
174
174
|
| 3 Design surface | UI tasks only: **Figma on or text-only** (super-ux's project-level choice, default on — check `docs/ux/foundation.md` → *Design tooling* before asking); is the Figma MCP connected; **and if it isn't — ship text-only, or stop here and connect it?** super-ux degrades to text-only on its own and never blocks, which means an unasked question here silently ships a UI feature with no mockups |
|
|
175
|
-
| 3 Design file | Figma on only: **exactly which
|
|
175
|
+
| 3 Design file | Figma on only: **exactly which files — one per surface (App, Web, ASO) — in which team/org** — the recorded ones, or a URL the operator gives, or *create one in a named team* with that creation explicitly authorized. A destination decided at drawing time is how a project ends up with three "design" files and no way to tell which is real. See *The design destination* below |
|
|
176
176
|
| 4–5 Dev | base branch; worktree/branch policy; is `main` off-limits; commit convention; task tracker |
|
|
177
177
|
| 5 Integration | how the branch lands (merge / PR + approver / "leave it unmerged"); parallel fan-out wanted (one worktree per implementer)? |
|
|
178
178
|
| 6 Tests | the test command; what "green" means here; known-red baseline; coverage expectation |
|
|
@@ -188,7 +188,7 @@ preconditions ("staging once lint and the full suite are green; production alway
|
|
|
188
188
|
asks"). Specific and recorded → it satisfies the stage-7 manual gate. Broader,
|
|
189
189
|
absent or ambiguous → stage 7 stops and asks.
|
|
190
190
|
|
|
191
|
-
## The design destination — one file, decided here, never invented later
|
|
191
|
+
## The design destination — one file per surface, decided here, never invented later
|
|
192
192
|
|
|
193
193
|
When the project designs in Figma, **the destination is a stage-0 decision, not a
|
|
194
194
|
stage-3 side effect.** Left to drawing time, the question "where do I put this?"
|
|
@@ -199,16 +199,22 @@ the team actually opens.
|
|
|
199
199
|
|
|
200
200
|
**Settle three things, in this order:**
|
|
201
201
|
|
|
202
|
-
1. **Is there already a file?** Read `docs/ux/foundation.md` →
|
|
203
|
-
first. A recorded, resolving file ends the question — record "use
|
|
204
|
-
file" and move on. Do not ask the operator something the project
|
|
202
|
+
1. **Is there already a file for this surface?** Read `docs/ux/foundation.md` →
|
|
203
|
+
*Design tooling* first. A recorded, resolving file ends the question — record "use
|
|
204
|
+
the recorded file" and move on. Do not ask the operator something the project
|
|
205
|
+
already answered.
|
|
205
206
|
2. **Which team / organization**, by name. A file URL identifies a file; it does not
|
|
206
207
|
say whose workspace it lives in, and a design that lands in someone's personal
|
|
207
208
|
drafts instead of the team space is invisible to everyone who needs it. When the
|
|
208
209
|
operator belongs to several teams, the choice is theirs and it gets written down —
|
|
209
210
|
`whoami` tells you which are available, it does not tell you which is right.
|
|
210
|
-
3. **Which file** — an existing URL the operator supplies, or **creation
|
|
211
|
-
named team, explicitly authorized.**
|
|
211
|
+
3. **Which file, per surface** — an existing URL the operator supplies, or **creation
|
|
212
|
+
in that named team, explicitly authorized.** The shape is **one file per surface** —
|
|
213
|
+
App, Web, ASO (store screenshots, icon, logo) — never one file holding every frame:
|
|
214
|
+
the surfaces have different owners, sizes and reviewers, and a single file is the
|
|
215
|
+
one nobody can hand to any of them. The record names each file with its surface,
|
|
216
|
+
and stage 3's gate checks every frame link against that set
|
|
217
|
+
(`scripts/visual_gate.py filekeys`).
|
|
212
218
|
|
|
213
219
|
**Creating a file in a shared workspace is outward and irreversible enough to need a
|
|
214
220
|
named target.** It follows the same floor as deploy authorization above: *"create
|
|
@@ -23,6 +23,7 @@ searches.
|
|
|
23
23
|
- Bookkeeping — the thing that makes detection mechanical
|
|
24
24
|
- Detection — any one of these trips the guard
|
|
25
25
|
- The review loop — a cap that measures rather than stops
|
|
26
|
+
- The re-render loop — a budget of one, two at most
|
|
26
27
|
- The break protocol
|
|
27
28
|
- When to stop and hand back
|
|
28
29
|
- Rationalizations
|
|
@@ -110,6 +111,28 @@ that was not.
|
|
|
110
111
|
`touch:` lines at the review stage. A round that finds nothing ends the loop by
|
|
111
112
|
definition and needs no counting.
|
|
112
113
|
|
|
114
|
+
## The re-render loop — a budget of one, two at most
|
|
115
|
+
|
|
116
|
+
A visual surface has a loop of its own: render, critique the render, render again. It
|
|
117
|
+
is the one loop here with a **budget rather than a cap**, because its gain is measured
|
|
118
|
+
to flatten fast — past the first or second refinement the change sits inside the noise
|
|
119
|
+
of the judge reading it, and a third round mostly swaps one defect for another. So:
|
|
120
|
+
|
|
121
|
+
- **One re-render per critique, two at the most.** After the second, the item still
|
|
122
|
+
failing is marked **`unresolved`** on the contact sheet with its triple and goes to the
|
|
123
|
+
person on their one pass — not into round three. Keep every rendered version; the last
|
|
124
|
+
is not automatically the best, and the sheet can show two side by side.
|
|
125
|
+
- **Only external, specific feedback starts a round**: a linter finding, a failed audit
|
|
126
|
+
item, a diff against the frame or the approved baseline, a *region → defect → fix*
|
|
127
|
+
triple. "Look again" or "make it better" is not feedback and starts nothing — a round
|
|
128
|
+
with no named defect is churn with a screenshot.
|
|
129
|
+
- **The person's rounds are counted, not remembered.** Each return of the contact sheet
|
|
130
|
+
with triples is a `review:` line in the run ledger ([`../templates/run.md`](../templates/run.md)),
|
|
131
|
+
and the sheet's `review_rounds` carries the same number
|
|
132
|
+
([`browser.md`](browser.md) → *The visual half*). Like the review cap above it is a
|
|
133
|
+
measurement — a class of defect that keeps reaching the person is a rule missing from
|
|
134
|
+
the machine checks, and the fix is that rule, not more attention.
|
|
135
|
+
|
|
113
136
|
## The break protocol
|
|
114
137
|
|
|
115
138
|
When the guard trips, **stop editing immediately**. Do not dispatch another fix, do
|
|
@@ -74,7 +74,7 @@ a row pointing outside the bundle is the defect this file exists to catch.
|
|
|
74
74
|
| What a spec must lock, the UX-track order, the module dossier | `references/spec.md` |
|
|
75
75
|
| The zero-context plan format, parallel groups, set equality | `references/planning.md` |
|
|
76
76
|
| The work graph: its fields, the verbs and their exit codes, and the three invariants a schema cannot state | `references/work-graph.md`, `scripts/graph.py`, `graph.schema.json` |
|
|
77
|
-
| How a node is CLOSED: three blind readings at escalating visibility, all three required
|
|
77
|
+
| How a node is CLOSED: three blind readings at escalating visibility, all three required — four on a flagship or product surface — and the round ledger the ceiling reads | `references/certification.md`, `agents/verifier-{unit,seam,product,visual}.md`, `scripts/graph.py certify` |
|
|
78
78
|
| Workspace isolation, the subagent loop, who may write the register | `references/build.md` |
|
|
79
79
|
| The review rubric, diff packages, the three verdicts | `references/review.md` |
|
|
80
80
|
| **False success** — the class, its known shapes and its two rules | `references/gates.md` |
|
|
@@ -33,9 +33,9 @@ Runs on **super-ux** — the one companion this pipeline recommends by name
|
|
|
33
33
|
give the install line and stop; don't improvise a half-chain.
|
|
34
34
|
|
|
35
35
|
0. **The design destination is already decided — read it, don't re-open it.** When
|
|
36
|
-
Figma is on, the stage-0 brief names the team/org and the
|
|
37
|
-
([`grill.md`](grill.md) → *The design destination*), and
|
|
38
|
-
`docs/ux/foundation.md` → *Design tooling* is the canonical record. Confirm
|
|
36
|
+
Figma is on, the stage-0 brief names the team/org and the files — one per surface
|
|
37
|
+
(App, Web, ASO) — ([`grill.md`](grill.md) → *The design destination*), and
|
|
38
|
+
`docs/ux/foundation.md` → *Design tooling* is the canonical record. Confirm each
|
|
39
39
|
recorded file **resolves** before any drawing. **Never create a file when a
|
|
40
40
|
recorded one resolves; if it doesn't resolve, stop and ask — never create a
|
|
41
41
|
replacement.** A creation happens at most once per project, in the team the
|
|
@@ -90,6 +90,30 @@ green over both.
|
|
|
90
90
|
boundary with Figma (tokens as variables, never raw values carried across). Not
|
|
91
91
|
through it: a purely structural change — what sits where is the UX track's — text,
|
|
92
92
|
a backend, an internal script.
|
|
93
|
+
- **The VISUAL track leaves a trace, and the gate reads the trace.** "The track ran" is
|
|
94
|
+
not checkable; a record is. The track writes a **director record** in the product's
|
|
95
|
+
repository, beside `docs/ux/`: `docs/design/<surface>/director-record.md`, a header
|
|
96
|
+
line `surface_class: <class>` and one `## <Field>` heading per field. Which fields are
|
|
97
|
+
owed is the brief's `surface_class` ([`stages.md`](stages.md) → stage 0, *The surface
|
|
98
|
+
class*):
|
|
99
|
+
|
|
100
|
+
| Class | Fields the record owes |
|
|
101
|
+
|---|---|
|
|
102
|
+
| `flagship` | Brief, Mode, Taste, References, Cast, Fork, Rubric, Critique, Markers, Alignment, Quality, Signature, Surfaces, Haptics, ADA, Open |
|
|
103
|
+
| `product` | Brief, Mode, References, Markers, Open |
|
|
104
|
+
| `ad` | Brief, Mode, References, Markers, ADA (the ad rubric profile and safe zones), Open |
|
|
105
|
+
| `internal` | none — the project linter is the floor |
|
|
106
|
+
|
|
107
|
+
What each field must say is sheleg-design's contract, and its validator checks it:
|
|
108
|
+
`npx sheleg-design-skill --check-record <file>`. The pipeline's gate is `python3
|
|
109
|
+
scripts/visual_gate.py record <file> --class <surface_class>`: it checks the headings
|
|
110
|
+
itself — present, and not a placeholder — and runs that validator where sheleg-design
|
|
111
|
+
is installed and new enough. Where it is not, the validator reads **NOT_RUN**, never
|
|
112
|
+
PASS, and the gate stands on the heading floor with that said. **The refusal is the
|
|
113
|
+
same file**: `## Mode` saying `declined` and why — *«без дизайна»*, *as is* — passes,
|
|
114
|
+
and a bare `declined` with no reason does not. The Rubric is written **before** any
|
|
115
|
+
direction is rendered; a rubric written after the render grades the render it already
|
|
116
|
+
liked.
|
|
93
117
|
- **Each track's refusal is a sentence, never a silence.** *"Без дизайна" / "as is"*
|
|
94
118
|
ends the visual track; *"без бренда" / "draft"* ends the copy track. Either one is
|
|
95
119
|
the operator's to make and costs nothing — but it is **recorded in the brief and
|