task-pipeline-skill 1.87.1 → 1.89.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (32) hide show
  1. package/CHANGELOG.md +116 -0
  2. package/README.md +1 -1
  3. package/SKILL-CARD.md +1 -1
  4. package/package.json +3 -3
  5. package/plugins/task-pipeline/.claude-plugin/plugin.json +1 -1
  6. package/plugins/task-pipeline/agents/verifier-product.md +3 -1
  7. package/plugins/task-pipeline/agents/verifier-visual.md +115 -0
  8. package/plugins/task-pipeline/agents/verifier.md +2 -1
  9. package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +4 -4
  10. package/plugins/task-pipeline/skills/task-pipeline/graph.schema.json +22 -1
  11. package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +4 -4
  12. package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +12 -0
  13. package/plugins/task-pipeline/skills/task-pipeline/references/audit.md +14 -8
  14. package/plugins/task-pipeline/skills/task-pipeline/references/browser.md +97 -3
  15. package/plugins/task-pipeline/skills/task-pipeline/references/certification.md +46 -5
  16. package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +28 -0
  17. package/plugins/task-pipeline/skills/task-pipeline/references/continuity.md +17 -0
  18. package/plugins/task-pipeline/skills/task-pipeline/references/conventions.md +3 -3
  19. package/plugins/task-pipeline/skills/task-pipeline/references/doctrine-map.md +1 -1
  20. package/plugins/task-pipeline/skills/task-pipeline/references/grill.md +15 -9
  21. package/plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md +23 -0
  22. package/plugins/task-pipeline/skills/task-pipeline/references/portability.md +1 -1
  23. package/plugins/task-pipeline/skills/task-pipeline/references/progress.md +5 -1
  24. package/plugins/task-pipeline/skills/task-pipeline/references/spec.md +27 -3
  25. package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +75 -13
  26. package/plugins/task-pipeline/skills/task-pipeline/references/work-graph.md +2 -2
  27. package/plugins/task-pipeline/skills/task-pipeline/scripts/graph.py +45 -16
  28. package/plugins/task-pipeline/skills/task-pipeline/scripts/stage_checkpoint.py +292 -0
  29. package/plugins/task-pipeline/skills/task-pipeline/scripts/visual_gate.py +589 -0
  30. package/plugins/task-pipeline/skills/task-pipeline/templates/brief.md +5 -2
  31. package/plugins/task-pipeline/skills/task-pipeline/templates/browser-claims.json +223 -1
  32. package/plugins/task-pipeline/skills/task-pipeline/templates/run.md +11 -1
@@ -28,6 +28,7 @@ which is how a run says *I checked the browser* and means *I ran the unit tests*
28
28
  - Sessions, and why an agent needs them
29
29
  - Reading a look vs gating on one
30
30
  - "Tested in a browser" is three different claims
31
+ - The visual half — pixels against intent, under a contract
31
32
  - Getting past a login, and past a backend
32
33
  - When the look finds something: debugging the spec that missed it
33
34
  - Evidence a reader can open
@@ -48,7 +49,9 @@ Three consequences the doctrine rests on:
48
49
 
49
50
  - **A look costs a page of text and no vision model.** This is why the pipeline can ask
50
51
  for one at three stages without the cost being an argument. A `screenshot` exists in
51
- both channels and *is* pixels — take one for a human to look at, not for you to read.
52
+ both channels and *is* pixels — in the functional look, take one for a human, not for
53
+ you to read. Reading pixels against the design's intent is a different check with its
54
+ own contract: *The visual half*, below.
52
55
  - **The ref is a fact about the page as rendered**, so `click e12` after a snapshot is
53
56
  deterministic in a way a coordinate never is.
54
57
  - **A ref that no longer resolves is a finding, not an error to retry past.** The element
@@ -125,7 +128,7 @@ commonest way a run reports a green it does not have.
125
128
 
126
129
  | | What it is | What it proves | Where it counts |
127
130
  |---|---|---|---|
128
- | **The look** | an agent driving a page: open, snapshot, console, network | that this surface renders, right now, and what the browser said while it did | the **look**, stage 6 — recommended, never a gate |
131
+ | **The look** | an agent driving a page: open, snapshot, console, network | that this surface renders, right now, and what the browser said while it did | the **functional look**, stage 6 — recommended, never a gate. Its **visual half** is a gate on a `flagship`, `product` or `ad` surface — *The visual half*, below |
129
132
  | **The spec suite** | `playwright test` — the **test runner** | that the assertions someone wrote still hold, on the paths someone thought to write | the **suite** half of the stage-6 gate, counted with every other test |
130
133
  | **The library** | `require('playwright')` — `chromium`/`firefox`/`webkit`, `devices`, `request`, `selectors` | whatever your own script asserts; it is an automation API, not a test framework | wherever the project already runs it |
131
134
 
@@ -155,6 +158,95 @@ in the claim's own state (the initial screenshot closes nothing about opened/err
155
158
  a suite PASS closes no look claim; a toggle owes its full cycle; no browser channel
156
159
  is NOT_RUN with the reason.
157
160
 
161
+ ## The visual half — pixels against intent, under a contract
162
+
163
+ Everything above reads the **accessibility tree**: it proves the surface renders and the
164
+ browser stayed quiet. It cannot say whether the surface looks like what was designed —
165
+ the wrong weight on a heading, a card that lost its spacing at 200 % text, a dark theme
166
+ that drops a border, a frame that drifted from the Figma it was built from. That is a
167
+ different question, and asking it of a snapshot is how a run reports *it looks right*
168
+ having read no pixels at all. The **functional look** and the **visual half** are two
169
+ checks; neither discharges the other.
170
+
171
+ **When it runs, and when it gates — by the brief's `surface_class`**
172
+ ([`stages.md`](stages.md) → stage 0, *The surface class*):
173
+
174
+ | Class | Visual half at stages 5–6 | Human pass at stage 10 |
175
+ |---|---|---|
176
+ | `flagship` | **gate** — full matrix, pairwise across every axis, judge items per the full rubric | **gate** — the approved contact sheet |
177
+ | `product` | **gate** — every state, the mandatory pairs; pairwise holes reported | **gate** — the approved contact sheet |
178
+ | `ad` | **gate** — the ad rubric profile and the safe zones | **gate** — the approved contact sheet |
179
+ | `internal` | recommended — the project linter is the floor; a sheet, if made, is checked for honesty | the functional look closes it |
180
+
181
+ It gates only where the stage-3 VISUAL track ran. A recorded refusal (*«без дизайна»*,
182
+ `Mode: declined`) turns the visual half into the functional look plus the linter, said in
183
+ the close-out — the same rule as every other refusal in the pipeline.
184
+
185
+ **The checks run cheapest first, and the person last:**
186
+
187
+ 1. **Deterministic.** The project linter — `python3 scripts/visual_gate.py lint <dir>`,
188
+ which runs sheleg-design's `--lint` where it is installed and answers **NOT_RUN (exit 3)**
189
+ where it is not — plus axe or Lighthouse, the token check, the type checker. An S1
190
+ finding blocks, whatever anyone says about the picture later.
191
+ 2. **Regression** against the approved baseline, where one exists (`toHaveScreenshot`, or
192
+ the platform's snapshot test).
193
+ 3. **The matrix.** One frame per `SCR-NN` state the screen map lists (default, loading,
194
+ empty, error, offline, long-content, keyboard-up, first-run — the ones that apply),
195
+ across viewport × theme × text size × locale: **pairwise coverage** — every value of one
196
+ axis meets every value of every other in some frame — plus the **mandatory pairs**
197
+ (dark × large text, RTL × narrow). Never the full cross product: every state at every
198
+ combination of every axis value is hundreds of frames, and nobody reads them. Seed the
199
+ worst-case data first (`break-ui`, [`companion-skills.md`](companion-skills.md) →
200
+ *Visual lanes*) so the frames show long names and empty lists, not the demo account.
201
+ 4. **Each frame carries its capture record** — revision, route, state, viewport, locale,
202
+ theme, motion, captured-at, source — and is disqualified, not passed, when it is blank,
203
+ stale, of the wrong route or state, or taken before the fonts rendered. This is
204
+ sheleg-design's visual-review contract; the pipeline only refuses a frame without it.
205
+ Where a Figma frame or an approved baseline exists, the frame is **diffed against it**,
206
+ and a failing diff is never a PASS the run writes: either the build drifted, or the
207
+ person approves a new baseline.
208
+ 5. **The rubric**, read by a judge that is not the agent that built the surface: a
209
+ checklist per task, never a single score; pairwise only against the approved
210
+ reference and in both orders; three samples, and disagreement is `uncertain`, which
211
+ goes to the person rather than to a coin. **A judge item (J) is `NOT_ASSESSED` until a
212
+ labelled set exists and the judge's agreement with it is measured** — a verdict from an
213
+ uncalibrated judge is a guess with a format. A gate item (G) is deterministic and the
214
+ judge never overrides its FAIL. Every FAIL is a triple: *region → defect → fix*.
215
+ 6. **One human pass** over the contact sheet. Approval makes its frames the next baseline.
216
+
217
+ **The contact sheet is the one surface the person reviews**, and its data is
218
+ `templates/browser-claims.json` — the same file, **not a second schema**. A look row
219
+ that carries `axes` is a frame: `axes{viewport,theme,text,locale}`, `capture{revision,
220
+ route,motion,captured_at,source}`, `figma_frame`, `baseline`, `diff`, `rubric[]`. The file
221
+ names its `surface`, `revision`, `review_rounds`, `approved_by` and `approved_at`. The
222
+ viewing page is a self-contained local HTML beside the frames — a full document, `<meta
223
+ charset="utf-8">`, no CDN. **Frames and the HTML stay out of git; the JSON goes in.**
224
+
225
+ ```bash
226
+ python3 scripts/visual_gate.py sheet design/review/contact-sheet.json \
227
+ --class product --artifact-root design/review --states SCR-01/default,SCR-01/empty
228
+ # stage 10 adds --require-approval
229
+ ```
230
+
231
+ Exit `0` PASS · `1` FAIL · `2` unreadable · `3` NOT_RUN — a frame that did not run makes
232
+ the whole verdict NOT_RUN, never PASS. **On a gated class NOT_RUN stops the stage and
233
+ asks**: connect a capture channel, or the operator accepts the surface as `unverified`,
234
+ recorded in the brief and named in the close-out. It is never reported around, which is
235
+ the failure `DEC-0004` kept the functional look ungated to avoid; `DEC-0006` gates the
236
+ visual half on these classes and keeps that exit open in words.
237
+
238
+ **Two rules keep the loop from becoming the work** ([`loop-guard.md`](loop-guard.md) →
239
+ *The re-render loop*): a re-render budget of **one, two at most**, after which a failing
240
+ item goes to the person as `unresolved`; and **only external, specific feedback** starts a
241
+ round — a linter line, a diff, an audit item, a triple — never "look again". Each return
242
+ from the person adds one to `review_rounds` and one `review:` line to the run ledger, so the
243
+ number of passes a surface took is measured, not remembered.
244
+
245
+ **Native surfaces are not web surfaces.** A web render styled as a phone is a mockup. A
246
+ native screen's frames come from a simulator or a device (XCUITest snapshots, Compose
247
+ screenshot tests) and its accessibility from the platform's own audit; with neither, the
248
+ native rows stay `NOT_RUN`, said in words.
249
+
158
250
  ## Getting past a login, and past a backend
159
251
 
160
252
  A surface behind auth is the usual reason a run skips the look. Both channels solve it,
@@ -277,7 +369,9 @@ On a CI box, headed is the failure you will spend an hour on.
277
369
  | The excuse | Why it fails |
278
370
  |---|---|
279
371
  | *"`playwright test` is green, the surface is checked."* | The suite asserts what someone wrote down. `DEC-0004`: it is the coverage half, never the look. |
280
- | *"I took a screenshot, so I looked."* | A screenshot is pixels you did not read. The look is `snapshot` + `console` + `requests`, and the verdict quotes them. |
372
+ | *"I took a screenshot, so I looked."* | A screenshot is pixels you did not read. The functional look is `snapshot` + `console` + `requests`, and the verdict quotes them; reading the pixels is the visual half, with a capture record per frame and a matrix behind it. |
373
+ | *"The snapshot is clean, so it looks right."* | The tree says the heading exists, not that it rendered at the right weight, in the dark theme, at 200 % text. That is the visual half's question, and on a flagship, product or ad surface it is a gate. |
374
+ | *"The judge said it looks great."* | An uncalibrated judge's J items are `NOT_ASSESSED`, and no judge overrides a deterministic FAIL. A verdict needs a checklist, a reference, both orders and three samples. |
281
375
  | *"The click failed, I'll find a better selector."* | A ref that stopped resolving **is the finding**. Re-snapshot and report what moved. |
282
376
  | *"The docs say the CLI has no `tracing`."* | A vendor page is a claim; `--help` is the tool. This file was written against `--help` **because** a page-derived claim shipped here and was wrong. |
283
377
  | *"The tool list is in the docs."* | The page listed tools this version does not ship, and omitted that tracing, video and PDF need `--caps`. Ask the server: 24 tools default, 42 with all caps. |
@@ -15,6 +15,7 @@ queue is — graph or plan — is what decides, not the mood of the closer.
15
15
 
16
16
  - Why one verifier is not enough, stated as the failure it produces
17
17
  - The three tiers
18
+ - The fourth reading — `visual`, on a flagship or product surface
18
19
  - Blind, and it is the whole design
19
20
  - A pass has to mean something, so two rules have teeth
20
21
  - The report, and where each field lands
@@ -60,7 +61,42 @@ that reads no code is not the soft one.
60
61
 
61
62
  Agents: [`../../../agents/verifier-unit.md`](../../../agents/verifier-unit.md),
62
63
  [`verifier-seam.md`](../../../agents/verifier-seam.md),
63
- [`verifier-product.md`](../../../agents/verifier-product.md).
64
+ [`verifier-product.md`](../../../agents/verifier-product.md) — and, on a visual surface,
65
+ [`verifier-visual.md`](../../../agents/verifier-visual.md), below.
66
+
67
+ ## The fourth reading — `visual`, on a flagship or product surface
68
+
69
+ The three tiers read code, what reaches it, and what the product says about it. **None
70
+ of them opens a picture**, and on a surface whose look is part of the requirement that
71
+ leaves a whole level unread: a node can pass unit, seam and product while its empty
72
+ state renders grey on grey at 200 % text, and every report is truthful about what its
73
+ tier saw.
74
+
75
+ | Tier | Subject | Characteristic finding |
76
+ |---|---|---|
77
+ | `visual` | the contact sheet, the director record, the project linter's output, the rubric | a frame that contradicts the record's intent; a gate item failing under a PASS; a hole in the state × axes matrix; a judge verdict from an uncalibrated judge |
78
+
79
+ **When it runs.** A node that builds a user-facing surface copies the brief's
80
+ `surface_class` ([`stages.md`](stages.md) → stage 0, *The surface class*). On
81
+ **`flagship` and `product`** the `visual` report is **required** — `graph.py certify`
82
+ refuses the round without it and names the tier. On `internal` and `ad` it is accepted
83
+ when given and counted like any other tier, and not demanded: an internal tool's floor
84
+ is the linter, and an ad's gate is its rubric profile on the contact sheet.
85
+
86
+ **It is the fourth blind reading, not a reviewer with a vision model.** Same eight-key
87
+ report, same `breaks`/`risk`, same blindness — it never sees the other three reports
88
+ and they never see it. Its own rules, because a judge of pixels fails in its own ways:
89
+
90
+ - **A checklist per task, never a score.** It answers the rubric's binary items against
91
+ the record and the sheet; "looks polished" is not an item.
92
+ - **Pairwise only against the approved reference, and in both orders**; three samples.
93
+ A verdict that flips with the order, or between samples, is **`uncertain`** and goes to
94
+ the person — it is neither a pass nor a `breaks`.
95
+ - **It never overrides a deterministic FAIL.** A gate item (G) or a linter S1 that
96
+ failed is a `breaks` whatever the picture looks like to it.
97
+ - **A judge item (J) is `NOT_ASSESSED` until a labelled set exists** and the judge's
98
+ agreement with it has been measured. It goes in `not_examined`, which reaches the
99
+ closing verdict as `not_verified` — the honest name for an opinion nobody calibrated.
64
100
 
65
101
  ## Blind, and it is the whole design
66
102
 
@@ -71,8 +107,10 @@ will paraphrase it back as product truth. The disagreement between blind reading
71
107
  the instrument, so `graph.py certify` refuses a report whose prose cites another
72
108
  tier's verdict.
73
109
 
74
- Dispatch all three in one message so they run concurrently. Give each the node id,
75
- its `serves`, and the diff — nothing else, and never another tier's output.
110
+ Dispatch all three in one message so they run concurrently — all four on a visual
111
+ node. Give each the node id, its `serves`, and the diff — the `visual` tier also the
112
+ paths of the contact sheet, the director record and the linter output — nothing else,
113
+ and never another tier's output.
76
114
 
77
115
  **The second axis, and it is the one an optimisation removes first: whoever produced
78
116
  the fix never grades it.** Tier blindness is horizontal — no tier reads another's
@@ -146,6 +184,9 @@ consumer refuses.
146
184
  # three reports in, one verdict out — exits 1 if any tier failed
147
185
  graph.py certify --node N-007 \
148
186
  --tier unit.json --tier seam.json --tier product.json
187
+ # a node with surface_class flagship or product: four, or the round is refused
188
+ graph.py certify --node N-008 \
189
+ --tier unit.json --tier seam.json --tier product.json --tier visual.json
149
190
 
150
191
  # unchanged, and still the only thing that moves the graph
151
192
  graph.py close --verdict .task-pipeline/verdict-N-007.json
@@ -176,8 +217,8 @@ has failed **every** round. A run spinning on one level needs the operator to se
176
217
 
177
218
  ## What this costs, said out loud
178
219
 
179
- Three agents per node instead of one. That is the price of the visibility, and it is
180
- paid per node rather than per run. The three are dispatched in parallel, so the
220
+ Three agents per node instead of one — four on a flagship or product surface. That is
221
+ the price of the visibility, and it is paid per node rather than per run. The three are dispatched in parallel, so the
181
222
  wall-clock cost is roughly one reading; the token cost is three. A node whose
182
223
  `check` is mechanical and whose blast radius is genuinely nil still pays it — and a
183
224
  tier with nothing to find says so in `scope` and `not_examined` rather than being
@@ -24,11 +24,15 @@ one must never look alike.
24
24
  > stop at the first that answers. The step stays **recommended and never a gate**: a gate
25
25
  > an environment cannot satisfy is one an agent learns to report around, and *verified by
26
26
  > reading the diff* already prices the absence honestly (`docs/DECISIONS.md`).
27
+ > **`DEC-0006`** scopes that to the functional look: the **visual half** is a gate on a
28
+ > `flagship`, `product` or `ad` surface, and its absent tool reads NOT_RUN and stops to
29
+ > ask rather than passing ([`browser.md`](browser.md) → *The visual half*).
27
30
 
28
31
  ## Contents
29
32
 
30
33
  - Built in — nothing to install
31
34
  - The matrix
35
+ - Visual lanes — tools, never entry points
32
36
  - Optional bridge — substituting an external skill set
33
37
  - Preflight (emit before stage 0)
34
38
  - Is this skill itself current?
@@ -76,6 +80,30 @@ one must never look alike.
76
80
  | ~~superpowers~~ | — | **Not a dependency.** Stages 2/4/5/6 run on the built-in doctrine above. See *Optional bridge* | — |
77
81
  | ~~grill-me / grilling~~ | — | **Not a dependency.** The stage-0 grill is built in (`references/grill.md`) | — |
78
82
 
83
+ ## Visual lanes — tools, never entry points
84
+
85
+ The visual half ([`browser.md`](browser.md) → *The visual half*) asks for checks no
86
+ companion above owns alone. These are the public tools that do each one. **Each is a
87
+ tool a stage reaches for, never an entry point and never a second route**: the stage
88
+ decides when, the tool does one job, and the close-out names which one it took. None is
89
+ required; an absent one is a check recorded as NOT_RUN with its reason.
90
+
91
+ | Tool | What it is for | Where in the stages |
92
+ |---|---|---|
93
+ | `break-ui` | seeds worst-case data — the longest name, the empty list, the 4-digit badge, the slow response — so frames show what a real account shows, not the demo | stage 6, **before** the contact-sheet screenshots |
94
+ | `review-animations`, `improve-animations` | reviews motion against its intent: timing, easing, interruption, reduced-motion fallback | stage 6, on a surface with motion; findings as triples |
95
+ | `mobile-native` | the mobile-web surface: viewport, safe areas, touch targets, keyboard-up states | stages 5–6, mobile-web frames of the matrix |
96
+ | `animate-expo` | React Native and Expo motion, built and reviewed on the platform's own primitives | stage 5, an RN or Expo surface |
97
+ | `webapp-testing`, `chrome-devtools` (`take_screenshot`, `lighthouse_audit`) | the frames of the matrix at each viewport and theme, and a Lighthouse pass as the deterministic floor | stage 6, the capture and the floor; stage 8 on a deployed target |
98
+ | `accessibility-review`, `a11y-debugging` | accessibility review of the rendered surface and debugging what it finds | stage 6, beside axe or Lighthouse |
99
+ | XCUITest `performAccessibilityAudit` | the iOS platform's own accessibility audit, run in UI tests | stage 6, a native iOS surface |
100
+ | Compose `enableAccessibilityChecks` | Android's accessibility checks in Compose UI tests | stage 6, a native Android surface |
101
+ | Playwright `toHaveScreenshot` | screenshot regression against the approved baseline, per state and theme | stage 6 (regression), and the nightly or release baseline after |
102
+ | axe-core | the deterministic accessibility floor of a web surface | stage 6, first in the cheap-first order |
103
+
104
+ A native screen is captured on a simulator or a device; a web render styled as a phone is
105
+ a mockup, and the native rows of the sheet stay NOT_RUN without one.
106
+
79
107
  ## Optional bridge — substituting an external skill set
80
108
 
81
109
  An operator who already runs an equivalent skill set may map it onto stages 2/4/5/6
@@ -20,6 +20,7 @@ almost no window left, loses the middle of it, and re-derives what it already di
20
20
  - The evidence rule
21
21
  - What happens at the signal
22
22
  - The flush is not a new document
23
+ - Part 3 — workflow memory at the stage boundary
23
24
  - Rationalizations
24
25
 
25
26
  ## The limit, before the capability
@@ -305,6 +306,22 @@ copy of the truth, it is written once, nobody updates it, and the next run reads
305
306
  it as current. The artifacts above are read by later stages anyway; making them
306
307
  right costs nothing extra and pays twice.
307
308
 
309
+ ## Part 3 — workflow memory at the stage boundary
310
+
311
+ The ledger survives a compaction. It does not survive a quota that ran out on another
312
+ account, or a run picked up by a different agent on another machine. When the host has
313
+ Project Observatory's memory tools, **every gate that returns also writes a workflow
314
+ checkpoint**, and the next executor continues from the last finished stage.
315
+
316
+ - **When.** After the `stage:` line is appended, every time a gate returns, whatever the verdict. A failed gate's checkpoint says `blocked` and keeps the stage open. Acceptance closes the workflow.
317
+ - **How.** `scripts/stage_checkpoint.py emit` prints the tool's arguments, built from the ledger alone: the topic, the verdicts, the next stage, the operator's constraints and key names. Pass them to `observatory_checkpoint_write`, or to `memory.checkpoint.write` under memory/0.1. Then hand the answer to `stage_checkpoint.py record --answer -`.
318
+ - **The constraints.** Give the brief's restrictive rules once, as `--constraint`. They are carried to every later boundary, and a successor reads them first.
319
+ - **Keys.** Pass keys by name, as `--credential PROJECT/ENV/NAME`, never as a value.
320
+ - **Without the tools.** `stage_checkpoint.py record --unavailable "<why>"` appends one `event: memory — unavailable` line, and the run continues exactly as before. Missing memory is a state the ledger names, not a failure.
321
+ - **A refusal.** `LeaseLost` means another executor holds the workflow now. The script drops the token. Read the workflow (`observatory_checkpoint_latest` with `stage_checkpoint.py state`'s id) before doing anything else. A stale writer stops after one refusal.
322
+ - **The lease token.** It lives in `.task-pipeline/memory.json` (mode 0600, git-ignored) and never in the ledger, a reply or a commit.
323
+ - **What this is not.** Not a second ledger and not the narrative. The checkpoint is the work's state for the next executor. The wiki and the ledger keep everything else.
324
+
308
325
  ## Rationalizations
309
326
 
310
327
  | Excuse | Reality |
@@ -103,9 +103,9 @@ closing a stage with an unread CI verdict.
103
103
  - Host self-update rules (module docs, runbooks, agent-self cards, etc.) — update
104
104
  in the same change. Fix dangling links.
105
105
  - **The design destination, on a project with no `docs/ux/`.** When the work uses
106
- Figma but super-ux isn't in play, there is no `foundation.md` to hold the file, so
107
- the brief is canonical — and a brief is per-run. Write the team and the file URL
108
- into the host's own docs (`CLAUDE.md`, or the README) in this change, so the next
106
+ Figma but super-ux isn't in play, there is no `foundation.md` to hold the files, so
107
+ the brief is canonical — and a brief is per-run. Write the team and each surface's
108
+ file URL into the host's own docs (`CLAUDE.md`, or the README) in this change, so the next
109
109
  run reads the destination instead of creating a second file
110
110
  ([`grill.md`](grill.md) → *The design destination*).
111
111
  - **The code graph:** [graphify](https://github.com/Graphify-Labs/graphify) —
@@ -30,7 +30,7 @@ UX track on a user-facing task.
30
30
  | 3 Spec | `references/spec.md` |
31
31
  | 4 Plan | `references/planning.md` |
32
32
  | the queue the loop walks | `references/work-graph.md` |
33
- | 5–8 · how a **work-graph node** is CLOSED — three blind readings at three distances, all three required (ceiling 3); a **prose-plan task** closes through `review.md` instead — one reviewer, five-round cap | `references/certification.md` |
33
+ | 5–8 · how a **work-graph node** is CLOSED — three blind readings at three distances, all three required, plus a fourth `visual` reading on a flagship or product surface (ceiling 3); a **prose-plan task** closes through `review.md` instead — one reviewer, five-round cap | `references/certification.md` |
34
34
  | 5 Build (worktree, subagents, fix loop) | `references/build.md` + `references/review.md` |
35
35
  | 5–6 TDD + suite gate | `references/tdd.md` |
36
36
  | 5, 6, 8 The browser — the look, the spec suite, and the difference | `references/browser.md` |
@@ -18,7 +18,7 @@ coming back to the operator.
18
18
  - Phase 2 — the gap check, then the loop
19
19
  - Domain awareness
20
20
  - The autonomy sweep
21
- - The design destination — one file, decided here, never invented later
21
+ - The design destination — one file per surface, decided here, never invented later
22
22
  - The REQ spine — the grill's other hard output
23
23
  - Output
24
24
 
@@ -170,9 +170,9 @@ explicit "stop and ask me here":
170
170
  | 0 Docs regime | where settled things live (the decision home — **one** per project, and an existing `docs/adr/` **is** it), who may write it, whether a lease mechanism is present or the run is `ungated`, the gate command and its ratchet floors, and whether this run may raise a floor ([`documentation.md`](documentation.md)) |
171
171
  | 1 Docs | external libs/APIs/SDKs in play; any private ones context7 can't resolve → where their docs live |
172
172
  | 2 Decompose | is this a platform (several capabilities/surfaces) or one module? if platform: deploy cadence — per module or once at the end |
173
- | 2–3 Spec | UI verdict (arms super-ux); any scenario-tracing waiver |
173
+ | 2–3 Spec | UI verdict (arms super-ux) **and the surface class** — `flagship`, `product`, `internal` or `ad`, which selects the visual gate profile ([`stages.md`](stages.md) → stage 0, *The surface class*); any scenario-tracing waiver |
174
174
  | 3 Design surface | UI tasks only: **Figma on or text-only** (super-ux's project-level choice, default on — check `docs/ux/foundation.md` → *Design tooling* before asking); is the Figma MCP connected; **and if it isn't — ship text-only, or stop here and connect it?** super-ux degrades to text-only on its own and never blocks, which means an unasked question here silently ships a UI feature with no mockups |
175
- | 3 Design file | Figma on only: **exactly which file, in which team/org** — the recorded one, or a URL the operator gives, or *create one in a named team* with that creation explicitly authorized. A destination decided at drawing time is how a project ends up with three "design" files and no way to tell which is real. See *The design destination* below |
175
+ | 3 Design file | Figma on only: **exactly which files — one per surface (App, Web, ASO) — in which team/org** — the recorded ones, or a URL the operator gives, or *create one in a named team* with that creation explicitly authorized. A destination decided at drawing time is how a project ends up with three "design" files and no way to tell which is real. See *The design destination* below |
176
176
  | 4–5 Dev | base branch; worktree/branch policy; is `main` off-limits; commit convention; task tracker |
177
177
  | 5 Integration | how the branch lands (merge / PR + approver / "leave it unmerged"); parallel fan-out wanted (one worktree per implementer)? |
178
178
  | 6 Tests | the test command; what "green" means here; known-red baseline; coverage expectation |
@@ -188,7 +188,7 @@ preconditions ("staging once lint and the full suite are green; production alway
188
188
  asks"). Specific and recorded → it satisfies the stage-7 manual gate. Broader,
189
189
  absent or ambiguous → stage 7 stops and asks.
190
190
 
191
- ## The design destination — one file, decided here, never invented later
191
+ ## The design destination — one file per surface, decided here, never invented later
192
192
 
193
193
  When the project designs in Figma, **the destination is a stage-0 decision, not a
194
194
  stage-3 side effect.** Left to drawing time, the question "where do I put this?"
@@ -199,16 +199,22 @@ the team actually opens.
199
199
 
200
200
  **Settle three things, in this order:**
201
201
 
202
- 1. **Is there already a file?** Read `docs/ux/foundation.md` → *Design tooling*
203
- first. A recorded, resolving file ends the question — record "use the recorded
204
- file" and move on. Do not ask the operator something the project already answered.
202
+ 1. **Is there already a file for this surface?** Read `docs/ux/foundation.md` →
203
+ *Design tooling* first. A recorded, resolving file ends the question — record "use
204
+ the recorded file" and move on. Do not ask the operator something the project
205
+ already answered.
205
206
  2. **Which team / organization**, by name. A file URL identifies a file; it does not
206
207
  say whose workspace it lives in, and a design that lands in someone's personal
207
208
  drafts instead of the team space is invisible to everyone who needs it. When the
208
209
  operator belongs to several teams, the choice is theirs and it gets written down —
209
210
  `whoami` tells you which are available, it does not tell you which is right.
210
- 3. **Which file** — an existing URL the operator supplies, or **creation in that
211
- named team, explicitly authorized.**
211
+ 3. **Which file, per surface** — an existing URL the operator supplies, or **creation
212
+ in that named team, explicitly authorized.** The shape is **one file per surface** —
213
+ App, Web, ASO (store screenshots, icon, logo) — never one file holding every frame:
214
+ the surfaces have different owners, sizes and reviewers, and a single file is the
215
+ one nobody can hand to any of them. The record names each file with its surface,
216
+ and stage 3's gate checks every frame link against that set
217
+ (`scripts/visual_gate.py filekeys`).
212
218
 
213
219
  **Creating a file in a shared workspace is outward and irreversible enough to need a
214
220
  named target.** It follows the same floor as deploy authorization above: *"create
@@ -23,6 +23,7 @@ searches.
23
23
  - Bookkeeping — the thing that makes detection mechanical
24
24
  - Detection — any one of these trips the guard
25
25
  - The review loop — a cap that measures rather than stops
26
+ - The re-render loop — a budget of one, two at most
26
27
  - The break protocol
27
28
  - When to stop and hand back
28
29
  - Rationalizations
@@ -110,6 +111,28 @@ that was not.
110
111
  `touch:` lines at the review stage. A round that finds nothing ends the loop by
111
112
  definition and needs no counting.
112
113
 
114
+ ## The re-render loop — a budget of one, two at most
115
+
116
+ A visual surface has a loop of its own: render, critique the render, render again. It
117
+ is the one loop here with a **budget rather than a cap**, because its gain is measured
118
+ to flatten fast — past the first or second refinement the change sits inside the noise
119
+ of the judge reading it, and a third round mostly swaps one defect for another. So:
120
+
121
+ - **One re-render per critique, two at the most.** After the second, the item still
122
+ failing is marked **`unresolved`** on the contact sheet with its triple and goes to the
123
+ person on their one pass — not into round three. Keep every rendered version; the last
124
+ is not automatically the best, and the sheet can show two side by side.
125
+ - **Only external, specific feedback starts a round**: a linter finding, a failed audit
126
+ item, a diff against the frame or the approved baseline, a *region → defect → fix*
127
+ triple. "Look again" or "make it better" is not feedback and starts nothing — a round
128
+ with no named defect is churn with a screenshot.
129
+ - **The person's rounds are counted, not remembered.** Each return of the contact sheet
130
+ with triples is a `review:` line in the run ledger ([`../templates/run.md`](../templates/run.md)),
131
+ and the sheet's `review_rounds` carries the same number
132
+ ([`browser.md`](browser.md) → *The visual half*). Like the review cap above it is a
133
+ measurement — a class of defect that keeps reaching the person is a rule missing from
134
+ the machine checks, and the fix is that rule, not more attention.
135
+
113
136
  ## The break protocol
114
137
 
115
138
  When the guard trips, **stop editing immediately**. Do not dispatch another fix, do
@@ -74,7 +74,7 @@ a row pointing outside the bundle is the defect this file exists to catch.
74
74
  | What a spec must lock, the UX-track order, the module dossier | `references/spec.md` |
75
75
  | The zero-context plan format, parallel groups, set equality | `references/planning.md` |
76
76
  | The work graph: its fields, the verbs and their exit codes, and the three invariants a schema cannot state | `references/work-graph.md`, `scripts/graph.py`, `graph.schema.json` |
77
- | How a node is CLOSED: three blind readings at escalating visibility, all three required, and the round ledger the ceiling reads | `references/certification.md`, `agents/verifier-{unit,seam,product}.md`, `scripts/graph.py certify` |
77
+ | How a node is CLOSED: three blind readings at escalating visibility, all three required — four on a flagship or product surface — and the round ledger the ceiling reads | `references/certification.md`, `agents/verifier-{unit,seam,product,visual}.md`, `scripts/graph.py certify` |
78
78
  | Workspace isolation, the subagent loop, who may write the register | `references/build.md` |
79
79
  | The review rubric, diff packages, the three verdicts | `references/review.md` |
80
80
  | **False success** — the class, its known shapes and its two rules | `references/gates.md` |
@@ -396,9 +396,13 @@ again, and it looks like enforcement while being a mirror.
396
396
  Three moments the rail cannot show, recorded by `hooks/run-lifecycle.sh` as
397
397
 
398
398
  ```
399
- event: <compact|session-end|subagent> — <detail> — <ISO-8601>
399
+ event: <compact|session-end|subagent|memory> — <detail> — <ISO-8601>
400
400
  ```
401
401
 
402
+ The fourth kind, `memory`, is appended by `scripts/stage_checkpoint.py` rather than
403
+ by the hook. It records a workflow checkpoint written at a stage boundary, a refusal,
404
+ or that memory was unavailable ([`continuity.md`](continuity.md) → *Part 3*).
405
+
402
406
  The rail reads none of them; `checkup` reads `session-end`, which is how an
403
407
  abandoned run stops being invisible. Before this the ledger simply stopped at
404
408
  whatever stage the session died on — and a stopped ledger is indistinguishable
@@ -33,9 +33,9 @@ Runs on **super-ux** — the one companion this pipeline recommends by name
33
33
  give the install line and stop; don't improvise a half-chain.
34
34
 
35
35
  0. **The design destination is already decided — read it, don't re-open it.** When
36
- Figma is on, the stage-0 brief names the team/org and the file
37
- ([`grill.md`](grill.md) → *The design destination*), and
38
- `docs/ux/foundation.md` → *Design tooling* is the canonical record. Confirm the
36
+ Figma is on, the stage-0 brief names the team/org and the files — one per surface
37
+ (App, Web, ASO) — ([`grill.md`](grill.md) → *The design destination*), and
38
+ `docs/ux/foundation.md` → *Design tooling* is the canonical record. Confirm each
39
39
  recorded file **resolves** before any drawing. **Never create a file when a
40
40
  recorded one resolves; if it doesn't resolve, stop and ask — never create a
41
41
  replacement.** A creation happens at most once per project, in the team the
@@ -90,6 +90,30 @@ green over both.
90
90
  boundary with Figma (tokens as variables, never raw values carried across). Not
91
91
  through it: a purely structural change — what sits where is the UX track's — text,
92
92
  a backend, an internal script.
93
+ - **The VISUAL track leaves a trace, and the gate reads the trace.** "The track ran" is
94
+ not checkable; a record is. The track writes a **director record** in the product's
95
+ repository, beside `docs/ux/`: `docs/design/<surface>/director-record.md`, a header
96
+ line `surface_class: <class>` and one `## <Field>` heading per field. Which fields are
97
+ owed is the brief's `surface_class` ([`stages.md`](stages.md) → stage 0, *The surface
98
+ class*):
99
+
100
+ | Class | Fields the record owes |
101
+ |---|---|
102
+ | `flagship` | Brief, Mode, Taste, References, Cast, Fork, Rubric, Critique, Markers, Alignment, Quality, Signature, Surfaces, Haptics, ADA, Open |
103
+ | `product` | Brief, Mode, References, Markers, Open |
104
+ | `ad` | Brief, Mode, References, Markers, ADA (the ad rubric profile and safe zones), Open |
105
+ | `internal` | none — the project linter is the floor |
106
+
107
+ What each field must say is sheleg-design's contract, and its validator checks it:
108
+ `npx sheleg-design-skill --check-record <file>`. The pipeline's gate is `python3
109
+ scripts/visual_gate.py record <file> --class <surface_class>`: it checks the headings
110
+ itself — present, and not a placeholder — and runs that validator where sheleg-design
111
+ is installed and new enough. Where it is not, the validator reads **NOT_RUN**, never
112
+ PASS, and the gate stands on the heading floor with that said. **The refusal is the
113
+ same file**: `## Mode` saying `declined` and why — *«без дизайна»*, *as is* — passes,
114
+ and a bare `declined` with no reason does not. The Rubric is written **before** any
115
+ direction is rendered; a rubric written after the render grades the render it already
116
+ liked.
93
117
  - **Each track's refusal is a sentence, never a silence.** *"Без дизайна" / "as is"*
94
118
  ends the visual track; *"без бренда" / "draft"* ends the copy track. Either one is
95
119
  the operator's to make and costs nothing — but it is **recorded in the brief and