jig-ui 0.8.2 → 0.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -4,19 +4,38 @@ The user invoked `{{command_prefix}}{{args_placeholder}}`. Treat
4
4
  Available subcommands: {{subcommand_list}}. If it is empty or is not one of
5
5
  these, say so, list them, and stop.
6
6
 
7
- Run the matching CLI command with `{{scripts_path}}`, passing the flags through
8
- unchanged. Then do the work below for that subcommand. Read the command's full
9
- output — findings are ordered by severity, not position, so `head`, `tail`,
10
- `grep` and `jq` drop the ones that matter.
7
+ There are two kinds of subcommand, and they are not run the same way.
8
+
9
+ **CLI-backed** — `install`, `update`, `init`, `check`, `explain`, `verdicts`, `gate`, `probe`. Run the
10
+ matching command with `{{scripts_path}}`, passing the flags through unchanged,
11
+ then do the work below for that subcommand. Read the command's full output —
12
+ findings are ordered by severity, not position, so `head`, `tail`, `grep` and
13
+ `jq` drop the ones that matter.
14
+
15
+ **Agent procedures** — `decide`, `spec`, `mockup`, `make`, `critique`. There is no binary. Do not try to run one:
16
+ the section below **is** the command. A CLI can check what these produce; it
17
+ cannot do their work, because the work is judgment and authorship.
18
+
19
+ `decide` runs once per project. The other four run for each page, feature or
20
+ functionality, one at a time — never the whole product at once:
21
+
22
+ ```text
23
+ decide project-wide decisions, and the reason for each once per project
24
+
25
+ spec what exactly is being built: its smallest useful version, at every screen size
26
+ mockup low-fidelity design of that spec, reviewed before any code
27
+ make high-fidelity: the actual page or feature, built from the spec & mockup
28
+ critique scrutinises what was built against the rules, its spec & mockup
29
+ ```
11
30
 
12
31
  ## init
13
32
 
14
33
  **Settle the surfaces before you run it.** Mode is the most consequential thing
15
- `init` writes and the thing it is worst at choosing: with `--yes` it takes
16
- `'/' → product` without reading the project at all, and `{{rules_path}}/01-modes.md`
17
- rule 1 then makes that config outrank your own reading of the project from then
18
- on. A default chosen in a second binds the project indefinitely, and density is
19
- expensive to reverse.
34
+ `init` writes and the thing it cannot choose: with `--yes` it has nobody to ask, so
35
+ it declares no surface mapping at all, and every mode-gated rule stays silent until
36
+ someone does. A mode written into `{{config_file}}` then outranks your own reading
37
+ of the project from then on (`{{rules_path}}/01-modes.md` rule 1), so it has to be
38
+ chosen deliberately — and density is expensive to reverse.
20
39
 
21
40
  You are the half that can fix this, because you are talking to someone who knows
22
41
  the answer and the CLI is not. So, first:
@@ -36,6 +55,11 @@ the answer and the CLI is not. So, first:
36
55
  Where there is genuinely one surface, say so and move on; the point is that the
37
56
  mode was chosen rather than defaulted into.
38
57
 
58
+ **Surfaces declared after `init` do not reach the page until `init` runs again.**
59
+ Writing a mode into `{{config_file}}` changes what agents read; the token files the
60
+ page imports still carry the mode `init` wrote. Run `{{scripts_path}} init` again after
61
+ any change to `surfaces`, and `check` will say when the two disagree.
62
+
39
63
  Afterwards, report what it detected, the brand colour it derived and where that
40
64
  came from, and the surfaces it used — the output says whether they came from
41
65
  `jig.config.json` or from the default.
@@ -55,23 +79,56 @@ its own attestation says `judgment=not-run` to make that explicit.
55
79
 
56
80
  Do the other half yourself:
57
81
 
58
- 1. Load `{{rules_path}}/00-anti-patterns.md` and `{{rules_path}}/05-copy.md`,
59
- plus the relevant section of `{{rules_path}}/03-patterns.md` for whatever the
60
- changed files build.
61
- 2. Apply the judgment rules to the same files the CLI just scanned.
82
+ 1. Load `rules.index.json` and filter it to the entries marked `pass: code`.
83
+ **That list is the work.** Do not decide from the filenames which rules are
84
+ likely to matter and read only those — the rule you would not have thought to
85
+ open is the one you are about to break. Read whatever files those ids live in.
86
+ 2. Return a verdict for **every id on the list**: `ok`, `finding`, or `n/a` with
87
+ one line of reason. `n/a, no form on this page` is a verdict; silence is not.
88
+ Leave the `pass: screen` rules alone: they need the rendered composition,
89
+ which is not in any file you have, and guessing at them here is how this
90
+ command and `critique` come to give one page two verdicts.
91
+ `{{command_prefix}}critique` owns those.
92
+
93
+ **This is self-review, and it is the weakest arm in the system.** You built
94
+ these files; you know what you meant, and what you meant is invisible in the
95
+ artifact. It runs here because it is better than nothing and it is cheap. It
96
+ does not replace `{{command_prefix}}critique`, whose arm C reads the same
97
+ rules against the same source with an agent that has never seen this
98
+ conversation. When the two disagree, the one that did not build the page is
99
+ the one to believe.
62
100
  3. Merge both halves into **one** report keyed by rule id, ordered by severity —
63
101
  not two lists. A reader should not have to know which half found what.
64
- 4. Run the self-check at the end of `{{rules_path}}/00-anti-patterns.md`.
102
+ 4. Run the self-check `L-04` at the end of `{{rules_path}}/00-anti-patterns.md`.
65
103
 
66
104
  Then emit the attestation with both halves filled in:
67
105
 
68
106
  ```text
69
- JIG_CHECK: version=<version> mode=<mode> mechanical=<pass|fail|skipped>:<n> judgment=<ran|skipped>
107
+ JIG_CHECK: version=<version> mode=<mode> mechanical=<pass|fail|skipped>:<n> warnings=<n> judgment=<ran|skipped>:<n> files=<n> styled=<n>
70
108
  ```
71
109
 
72
- Take `mechanical=` from the CLI's own line. If `check` could not run, that is
110
+ Take `mechanical=`, `warnings=`, `files=` and `styled=` from the CLI's own line.
111
+ `warnings=` counts the findings `check` reported as warnings. `mechanical=pass`
112
+ means no errors, and nothing more: a page that is not usable on a phone can
113
+ carry `pass:0` with several warnings, because the mobile detectors warn rather
114
+ than fail CI. A record that says `pass` with warnings above zero is not a clean
115
+ page — say what the warnings were.
116
+
117
+ If `check` could not run, that is
73
118
  `mechanical=skipped:0` — never `pass`, which would report a clean result for a
74
- check that inspected nothing. Report `judgment=ran` only if you did step 2.
119
+ check that inspected nothing.
120
+
121
+ `judgment=`'s `<n>` is **the number of ids you returned a verdict for**, not the
122
+ number of findings. It exists because `ran` on its own cannot be checked: an
123
+ agent that read four rules and an agent that walked all of them both wrote
124
+ `ran`. A number that is far below the size of the `pass: code` list is a
125
+ sampled review, and it says so without anyone having to ask.
126
+
127
+ `judgment=ran` here means the `code` rules ran. It does not mean the page was
128
+ reviewed — nothing in this command looks at a rendered page, so a screen whose
129
+ stylesheet never loaded passes every part of it. Say so if the user is treating
130
+ a clean `check` as a finished review, and point them at
131
+ `{{command_prefix}}critique`.
75
132
 
76
133
  ## install
77
134
 
@@ -79,8 +136,23 @@ Report where the skill landed and at which scope. If it warned that a global
79
136
  install already exists, do not work around it — that warning is the system
80
137
  refusing to leave two contradictory skills for one harness.
81
138
 
139
+ **The Stop hook is the user's to accept, not yours.** In Claude Code at project
140
+ scope, `--hook` adds an entry to `.claude/settings.json` that runs `{{scripts_path}}
141
+ gate` when an agent tries to finish, and holds it there while the files it changed
142
+ fail `check` or the step it just ran is unfinished. Run without the flag it is not
143
+ added, and `--no-hook` says so outright. Do not pass `--hook` on the user's behalf:
144
+ tell them it exists, say plainly that it can stop an agent finishing, and let them
145
+ answer. A hook you chose for them is a change to their editor they did not make.
146
+
82
147
  ## explain
83
148
 
149
+ `--layer` is the entry point for an agent that knows what it is trying to do but
150
+ not the rule number. `explain --layer` names the six and the question each
151
+ answers; `explain --layer layout` lists that one. The layers are a view, not a
152
+ location — nothing was renumbered to fit one, so `P-02 Button` sits in
153
+ `components` and `P-04 Form` in `patterns` while both keep the `P-` address they
154
+ have always had.
155
+
84
156
  Print the CLI's output as it stands. It is already the rule's full text — do not
85
157
  summarise it, and do not paraphrase the correction into your own words: the
86
158
  wording is the rule.
@@ -118,3 +190,966 @@ locally and are the user's; never re-apply Jig's version over them.
118
190
 
119
191
  Everything in `{{rules_path}}/` is yours to read. Cite rules by id, and cite any
120
192
  rule you deliberately break with the reason, in one line.
193
+
194
+ ## spec
195
+
196
+ Say exactly what is being built — a page, a feature, or a piece of
197
+ functionality — before any of it is designed. This command writes no markup and no
198
+ CSS.
199
+
200
+ **Start with what it is for, not how it is laid out.** "A page where the user
201
+ creates an invoice" is a spec. "A dashboard with a sidebar" is a layout with
202
+ nothing in it yet: a sidebar, a container width and a navigation structure are
203
+ guesses until there is something real to put in them.
204
+
205
+ ### Refuse rather than guess
206
+
207
+ Stop and say which of these is missing, rather than proceeding:
208
+
209
+ - `{{config_file}}` does not exist → run `init` first. Without it there is no
210
+ mode and no token layer, so there is nothing for a spec to be specific to.
211
+ - No `DECISIONS.md` with substance → run `{{command_prefix}}decide` first. A spec
212
+ implements the project's decisions, and cannot check itself against ones that
213
+ do not exist.
214
+ - The user named nothing → ask which page, feature or functionality.
215
+ `{{command_prefix}}spec` with no argument is not a request to invent one.
216
+
217
+ Read `DECISIONS.md` before asking anything. Do not re-ask what it already records;
218
+ cite it instead. If the screen touches an item in its **Unresolved** section, ask the user
219
+ about it by name; do not settle it in the spec.
220
+
221
+ ### 1. Ask, two or three questions at a time
222
+
223
+ **This is a required interaction, not a suggestion.** Ask two or three questions,
224
+ then **wait**. Do not send a questionnaire. Do not write the spec and present it
225
+ for approval on your first response — a spec synthesised from a one-line prompt
226
+ and handed over for a yes is the failure this command exists to prevent, and it
227
+ looks exactly like the command working.
228
+
229
+ Ask about **structure**, not appearance. Colour, type scale and density were
230
+ settled at `init` and live in `{{config_file}}` and the token layer; re-deciding
231
+ them per screen creates a second source of truth for the values Jig exists to
232
+ centralise. If the user raises appearance, record it and move on.
233
+
234
+ - Round 1 — what is it for? What does the user need to do there ("create an
235
+ invoice")? Who arrives, and what is the single most important thing they must be
236
+ able to see or do?
237
+ - **Then scope it — be a pessimist.** Of everything this could include,
238
+ what is the smallest version that is useful on its own? Ask that plainly and
239
+ expect to cut. "Comments with attachments, mentions, reactions, editing,
240
+ threading and notifications" is a wishlist; "text and a submit button" might be
241
+ V1. What is cut goes in `later:`, by name — deferred on the record, not
242
+ forgotten. Everything after this designs V1 only.
243
+ - Round 2 — what content and states does it carry? Empty, loading, error,
244
+ first-run, and the realistic range (0 items, 5, 500)?
245
+ - Round 3 — only what is still unresolved: entry points, where it leads, what
246
+ must not happen here.
247
+
248
+ **Ask about the phone by name.** What does the reader see first on a phone, and
249
+ what moves, stacks, or goes behind a control? If the conversation only ever
250
+ describes the wide screen, the phone composition will be derived from it rather
251
+ than designed — and a derived phone composition is a squeezed desktop.
252
+
253
+ ### 2. Run `L-01`
254
+
255
+ Read `L-01 · Layout method` in `{{rules_path}}/03-patterns.md` and run its five
256
+ steps against the answers. The spec's structural half **is** that output — group,
257
+ order by importance, space from the inside out, align to a grid, and state the
258
+ hierarchy plainly enough that the squint question can be answered against it.
259
+
260
+ ### 3. Write it down — one composition per screen size
261
+
262
+ `.jig/specs/<surface-slug>.spec.md`. Frontmatter carries what can be compared to
263
+ a built page; prose carries the reasoning that no check can read.
264
+
265
+ **Write a whole composition for every size, starting with the phone.** A spec
266
+ with one set of regions and a note about what changes when it is narrow is one
267
+ spec doing the work of three. It is always the wide composition that gets
268
+ designed and the narrow one that gets derived, which is how a page arrives on a
269
+ phone as a squeezed desktop.
270
+
271
+ ```yaml
272
+ ---
273
+ feature: sign in to an existing account # the task, not the screen
274
+ surface: login page
275
+ mode: product # from {{config_file}} — state it, never leave it implied
276
+ confirmed: false # only the user sets this to true; see below
277
+ sizes: # phone first. Each size is a whole composition, not a diff.
278
+ phone: # judged at 360px
279
+ regions: # in order, top to bottom
280
+ - brand: wordmark only
281
+ - form: the single primary task
282
+ - recovery: secondary, below the action
283
+ hierarchy: [form, brand, recovery] # most important first
284
+ grouping: { form: proximity, recovery: continuity }
285
+ grid: { columns: 4 }
286
+ nav: none on this screen
287
+ tablet: # judged at 768px
288
+ same-as: phone
289
+ why: "one centred form; a second column would hold nothing"
290
+ desktop: # judged at 1280px
291
+ regions:
292
+ - brand: left column — wordmark and one line on what the product does
293
+ - form: right column, the single primary task
294
+ - recovery: below the action, in the form column
295
+ hierarchy: [form, brand, recovery]
296
+ grouping: { form: proximity, brand: common region }
297
+ grid: { columns: 12, form: 5, brand: 6 }
298
+ nav: none on this screen
299
+ states: [default, loading, error, success]
300
+ decisions: [The Engineer Reads First] # every DECISIONS.md entry this screen implements, by name
301
+ later: [social sign-in, remember this device] # cut from V1, by name — the next specs start here
302
+ mockup: pending # `mockup` sets approved or skipped, from the user's own words
303
+ mockup_at: # where the approved drawing is: a .jig/mockups path, or a Figma or Stitch link
304
+ deviations: [] # `make` writes here; `spec` leaves it empty
305
+ ---
306
+ ```
307
+
308
+ - **`phone` is written first, and in full.** It is the most common screen, and the
309
+ one a derived composition fails on worst.
310
+ - **`tablet` and `desktop` are written in full too.** `same-as: <size>` is allowed
311
+ when nothing changes, but only with `why:`. It is a claim that the composition
312
+ holds at that width — `critique` checks it on a render — not a way to leave the
313
+ size out.
314
+ - **The widths are where a composition is judged, not where breakpoints must
315
+ fall.** The CSS breaks wherever the content needs it, and there may be more
316
+ breakpoints than sizes. A navigation row, for example, breaks at the width where
317
+ its labels fit.
318
+ - **No shell before there are features.** If the product has no navigation yet,
319
+ `nav:` says `none yet` — it is decided once there is more than one feature to
320
+ move between, not invented for the first one.
321
+ - **`nav:` at every size, on any screen with navigation — decided by `P-14`'s table
322
+ at that width, not copied.** It is the field most likely to be written once and
323
+ copied, and most likely to be genuinely different at every size on a page that
324
+ works. Count the destinations and run the table for each size: five links fit on
325
+ one row at 768px and 1280px, so there the value is the links in a row, and the
326
+ Menu button exists only at the widths where they do not fit. In a live run all
327
+ four specs wrote "menu button, top right" at phone, tablet **and** desktop, citing
328
+ a `DECISIONS.md` entry that said only where the button sits. Every page then hid
329
+ five links behind a menu on a 1280px screen, and every later step faithfully
330
+ built and approved it.
331
+
332
+ Every frontmatter field must be answerable by looking at the built page at that
333
+ size. If a field cannot be, it is prose — put it below the frontmatter, where it
334
+ belongs. "Feels focused" is prose. "Three regions, form first" is a field.
335
+
336
+ Add fields the surface needs (`form: { steps, validation }` for a form). Do not
337
+ add fields nothing can check.
338
+
339
+ ### 3b. Check it against the decisions before anyone confirms it
340
+
341
+ A live run produced a spec that ranked the price above the plan limits, on a
342
+ product whose confirmed `DECISIONS.md` said the limits are the headline and the
343
+ price never dominates. The agent had written that decision itself, one command
344
+ earlier. Nothing in the chain would have caught it: `critique` compares the page
345
+ to the spec, never the spec to the decisions.
346
+
347
+ - List every `DECISIONS.md` entry the screen implements in `decisions:`, by name.
348
+ - Check **every size's** fields against each one. A field that contradicts a
349
+ decision is fixed now, before confirmation. It is not a deviation — deviations
350
+ record what building revealed, not what writing got wrong.
351
+ - **A menu button where the links fit is refused.** Before confirmation, read `nav:`
352
+ at every size. If it says a menu button at a width where `P-14`'s table says the
353
+ destinations fit, fix it now. A decision about where the menu button sits is
354
+ about position, not about whether a button exists at that width — it does not
355
+ override `P-14`.
356
+ - **Every reference must resolve.** A decision is cited by its name; a rule by an
357
+ id that `{{command_prefix}}explain` returns. Where nothing decides something,
358
+ write `unspecified — make chooses one it can defend, and records it`. Never write
359
+ a sentence that implies a rule exists when it does not. The same run wrote
360
+ "FAQ arranged per DECISIONS.md rules" where no such rule existed — and the next
361
+ agent would have gone looking for it, found nothing, and invented one.
362
+ - **"The design system decides" is not a value.** When the user defers — "you decide",
363
+ "whatever Jig says", "leave it to the system" — that is an instruction to look the
364
+ answer up, not a phrase to copy into the spec. Search in this order: the `P-`
365
+ pattern for the component (`{{command_prefix}}explain <component>`), `L-01`'s five
366
+ steps, then the rules. Write the composition that lookup gives — "stacked, Pro
367
+ first", "Menu button, links in a vertical list" — and cite the id it came from.
368
+ Three of six specs in a live run wrote "arrangement decided by design system" into
369
+ the frontmatter, and the build that read it had nothing to build from.
370
+ - **Before asking for confirmation, read every field back and refuse any value that
371
+ names nobody.** "Decided by the design system", "as appropriate", "per best
372
+ practice", "TBD", "standard layout" each hand the decision to no one. Replace each
373
+ with the looked-up composition. Only when the lookup genuinely finds nothing does
374
+ the field say `unspecified — make chooses one it can defend, and records it`; and
375
+ say in the confirmation message which fields those are, so the user can decide
376
+ them instead.
377
+
378
+ ### 4. Get it confirmed
379
+
380
+ Show the spec and ask the user to confirm it. Then **stop**.
381
+
382
+ `confirmed: true` records that **the user said so in their own response**. You
383
+ may not set it on their behalf, and you may not treat your own summary,
384
+ `{{config_file}}`, or the absence of objection as confirmation. An unconfirmed
385
+ spec is not a weaker spec — it is not a spec, and `make` will refuse it.
386
+
387
+ If the user changes something, revise and ask again. The run is incomplete until
388
+ they confirm.
389
+
390
+ ## mockup
391
+
392
+ **Low-fidelity design.** Draw the confirmed spec in grayscale, at every size it
393
+ names, and get it looked at **before** any code is written. Once approved, the
394
+ drawing is what `make` builds to and what `critique` checks against, alongside the
395
+ spec. The spec stays the durable record, and everything the review changes goes
396
+ back into it first, so the two never disagree.
397
+
398
+ It draws **what one spec defines** — its V1 — and nothing deferred to `later:`.
399
+ Never the whole product.
400
+
401
+ ### Refuse rather than guess
402
+
403
+ - No `.jig/specs/<surface>.spec.md` → run `{{command_prefix}}spec` first.
404
+ - `confirmed: false` → nobody has agreed to what you would be drawing. Ask, and stop.
405
+ - The request is "mock up the app" or "the whole site" → say that a mockup draws
406
+ one specified page, feature or functionality, and offer to run `spec` for it.
407
+
408
+ ### 1. Ask the user how they want it drawn
409
+
410
+ **Ask every time, before drawing anything.** Three ways, and the choice is the
411
+ user's:
412
+
413
+ 1. **HTML** — you draw it as `.jig/mockups/<surface-slug>.html`. Needs nothing
414
+ connected. `.jig/` is outside what `check` scans and outside what ships, which
415
+ is where a working drawing belongs.
416
+ 2. **Figma** — drawn on a Figma canvas through Figma's MCP server.
417
+ 3. **Google Stitch** — generated through Stitch's MCP server.
418
+
419
+ Jig ships neither tool and writes nothing for either. They are MCP servers the user
420
+ connects to their own agent, and the user decides whether to.
421
+
422
+ **If they choose Figma or Stitch, check that its server is among your tools.** If it
423
+ is not, tell them plainly: that tool's MCP server is not connected to this agent;
424
+ they need to connect it — Figma's through its official MCP server, Stitch's with a
425
+ Stitch API key — and some agents only see a newly added server after a restart.
426
+ Then **stop and wait**. When they say it is connected, check your tools again
427
+ before drawing. Do not switch to HTML on their behalf; offer it, and let them
428
+ choose.
429
+
430
+ - **Drawing in Figma.** One page named for the feature, one frame per size at the
431
+ spec's widths — 360, 768, 1280 — phone first. Grey fills and strokes only; no
432
+ library components, styles or variables from the product's design system, for
433
+ the same reason the HTML carries no tokens.
434
+ - **Drawing in Stitch.** Ask for a low-fidelity grayscale wireframe of this one
435
+ feature, at each size, with the spec's regions in order and its real labels.
436
+ Stitch leans towards finished, coloured screens: if what comes back is polished
437
+ — brand colour, imagery, decoration — it is not a mockup. Ask Stitch again for a
438
+ wireframe. If it still will not produce one, tell the user and ask whether to
439
+ switch to HTML. Do not review a finished-looking screen as though it were one.
440
+
441
+ **One screen per size, one request each:** `deviceType` `MOBILE` for 360, `TABLET`
442
+ for 768, `DESKTOP` for 1280. A live run asked once per project and got phone
443
+ screens only.
444
+
445
+ **A timeout is not a failure.** Generating a screen takes minutes, longer than the
446
+ call waits, so the call often times out while Stitch keeps working and saves the
447
+ screen afterwards. In a live run all three agents took the timeout as "Stitch made
448
+ nothing", retried — which made a duplicate — and switched tools, while their
449
+ screens were sitting in the project. So, when the call times out or drops:
450
+ 1. **Do not call it again.** A second request is a second screen.
451
+ 2. List the project's screens every 30 seconds or so, for up to 10 minutes. The
452
+ list of projects does not show screens; only listing one project's screens
453
+ does.
454
+ 3. When the screen appears, carry on with it. If the project already holds a
455
+ screen for that size from an earlier attempt, use it rather than generating
456
+ another.
457
+ 4. Only when 10 minutes pass with nothing, tell the user, and ask whether to try
458
+ again or switch to HTML or Figma. Do not switch on their behalf.
459
+
460
+ Before review, confirm there is a screen for **each** size, not only the phone.
461
+ - **An existing design the user already has** for this feature, in either tool, can
462
+ stand in for drawing one. Read it through the server, and compare its structure
463
+ to the spec at every size. It will usually be high fidelity, so say plainly that
464
+ this step reviews structure only, and do not take up its colours or type as
465
+ decisions — those belong to the tokens and `DECISIONS.md`.
466
+
467
+ The rules below apply whichever of the three it is. They are written for HTML; in
468
+ Figma or Stitch they apply to the frames.
469
+
470
+ ### 2. Draw it — structure only
471
+
472
+ **Low fidelity means low detail.** The purpose is to judge structure before
473
+ appearance, and appearance wins every time it is allowed in: a reviewer shown
474
+ finished colour and type stops evaluating the layout and starts reacting to the
475
+ look. So the drawing contains:
476
+
477
+ - **Grayscale only.** No brand colour, no accent, no status colours — unless colour
478
+ *is* the feature (a colour picker, a status legend), and then only there.
479
+ - **No tokens and no project CSS.** Use the wireframe stylesheet below, pasted into
480
+ the file's `<style>`. The mockup must not look like the product; if it does, it
481
+ has become a design review.
482
+ - **No icons, imagery, shadows or decoration.** An image is a labelled box that says
483
+ what goes there. An icon is its name in brackets: `[search]`.
484
+ - **Real words, not lorem ipsum.** Labels, headings and button text are part of the
485
+ structure — "Create invoice" versus "Submit" is a decision worth reviewing.
486
+ - **Every size the spec names, as its own frame** — phone, then tablet, then desktop,
487
+ stacked top to bottom, each drawn with its own markup. Size is structure, not
488
+ detail: a phone composition that is only the desktop one narrower is the failure
489
+ specs are written per size to prevent. A `same-as:` size is drawn anyway, so the
490
+ claim can be seen to hold.
491
+ - **Navigation at every size, exactly as the spec's `nav:` says for that size.** On
492
+ the phone that usually means the Menu button, drawn where `DECISIONS.md` puts it —
493
+ and, because an open menu changes the structure, a second phone frame labelled
494
+ `phone — menu open` showing the links as a vertical list and the button reading as
495
+ close (`P-14`). A phone frame with no way to the other pages is not a simpler
496
+ drawing; it is a missing region. A live run approved exactly that in Figma, and the
497
+ build copied the gap.
498
+ - **The primary state, plus any state that changes the structure.** An empty state
499
+ that replaces a table, or an error that moves a form, is drawn. The rest stay
500
+ listed in the spec. Do not draw what building will reveal: 2,000 contacts, a
501
+ thirty-character name, a slow network — unless the feature is about exactly that.
502
+
503
+ ```html
504
+ <style>
505
+ /* Jig wireframe. Structure only: do not add colour, fonts, shadows or images. */
506
+ *, *::before, *::after { box-sizing: border-box; }
507
+ body { margin: 0; padding: 24px; font: 16px/1.5 system-ui, sans-serif; color: #222; background: #f4f4f4; }
508
+ .size { margin: 0 0 8px; font-size: 13px; color: #666; }
509
+ .frame { background: #fff; border: 1px solid #999; margin: 0 0 40px; padding: 16px; overflow: hidden; }
510
+ .frame[data-size="phone"] { width: 360px; }
511
+ .frame[data-size="tablet"] { width: 768px; }
512
+ .frame[data-size="desktop"] { width: 1280px; }
513
+ .region { border: 1px dashed #aaa; padding: 12px; margin: 0 0 12px; }
514
+ .region > .name { display: block; font-size: 12px; color: #777; margin: 0 0 8px; }
515
+ .media { background: #ddd; color: #555; display: grid; place-items: center; min-height: 120px; }
516
+ .field { display: block; width: 100%; min-height: 40px; border: 1px solid #888; background: #fff; }
517
+ button { font: inherit; padding: 8px 16px; border: 1px solid #555; background: #fff; }
518
+ button.primary { background: #444; color: #fff; border-color: #444; }
519
+ h1 { font-size: 28px; margin: 0 0 12px; }
520
+ h2 { font-size: 22px; margin: 0 0 8px; }
521
+ h3 { font-size: 18px; margin: 0 0 8px; }
522
+ .dims { font-variant-numeric: tabular-nums; color: #555; }
523
+ </style>
524
+ <script>
525
+ // Jig wireframe. Labels every region with its measured size, e.g.
526
+ // "navigation · 326 × 48". Measured from the drawing, never typed by hand.
527
+ function jigDims() {
528
+ document.querySelectorAll('.region > .name').forEach((label) => {
529
+ const box = label.parentElement.getBoundingClientRect();
530
+ let dims = label.querySelector('.dims');
531
+ if (!dims) { dims = document.createElement('span'); dims.className = 'dims'; label.append(dims); }
532
+ dims.textContent = ` · ${Math.round(box.width)} × ${Math.round(box.height)}`;
533
+ });
534
+ }
535
+ addEventListener('load', jigDims);
536
+ addEventListener('resize', jigDims);
537
+ </script>
538
+ ```
539
+
540
+ Each frame is a `<section class="frame" data-size="phone">`, preceded by a
541
+ `<p class="size">` naming the size and its width. **Every region is labelled** —
542
+ navigation, header, each group, the primary action — with the name the spec uses,
543
+ in a `.name` label, so the drawing and the spec can be read against each other.
544
+
545
+ **Every label carries the region's dimensions in px**, width × height at that frame's
546
+ width: `navigation · 326 × 48`. In HTML the script above writes them; do not type
547
+ them yourself, because a number that was not measured is a guess the reviewer will
548
+ trust. In Figma, put the frame's own width and height in the layer's label. In
549
+ Stitch, ask for the labels in the prompt, and where the screen does not show them,
550
+ say so in review rather than supplying numbers.
551
+
552
+ **The dimensions describe the drawing; they are not the build's values.** A
553
+ navigation 326px wide at 360 is the phone width minus the frame's padding and border, not a width to
554
+ hard-code — `make` still sizes with tokens and fluid widths (`D-111`). What they are
555
+ for is review: is the header taking a sixth of the phone? Is a tap target under
556
+ `--size-touch-target`? Is the hero pushing the plans below the first screen?
557
+
558
+ ### 3. Put it in front of the user
559
+
560
+ **Check the drawing against the spec before anyone sees it.** For each size, go
561
+ down the spec's `regions:` and its `nav:` and find each one in that frame. A region
562
+ the spec names and the frame lacks is drawn now, not raised in review. This matters
563
+ most on the phone, where a region dropped for space is easy to miss and the reviewer
564
+ is looking at what is there, not at what is not.
565
+
566
+ Tell the user where the drawing is — the file path, or the Figma or Stitch link —
567
+ and how to open it. If the harness can render —
568
+ a browse skill, Playwright, a headless browser — render each frame and show them the
569
+ images; a drawing nobody looked at has not been reviewed.
570
+
571
+ Then ask **structural** questions, not whether they like it: is the most important
572
+ thing the first thing seen at each size? Does anything belong in a different group?
573
+ Is anything here that V1 does not need?
574
+
575
+ ### 4. Every change goes into the spec first
576
+
577
+ When the review moves, adds, cuts or regroups anything, **edit the spec**, then
578
+ redraw from it. Never adjust the drawing alone. A mockup changed in review and not in
579
+ the spec has quietly become the real specification, and `make` would build from a
580
+ document that no longer describes what was approved.
581
+
582
+ A request for more — "add reactions while you're at it" — goes to `later:`, unless
583
+ the user explicitly moves it into V1. Then the spec's scope changes, and it is
584
+ confirmed again.
585
+
586
+ ### 5. Record the outcome in the user's words
587
+
588
+ - The user approves → set `mockup: approved` in the spec, and `confirmed: true` if
589
+ review changed it. Record where the approved drawing is in `mockup_at:` — the
590
+ file path, or the Figma or Stitch link. Only their own response counts; your summary of the drawing,
591
+ or no objection, does not.
592
+ - The user says to skip it → `mockup: skipped — <their reason>`. For a change small
593
+ enough that drawing it costs more than building it, that is a reasonable call; it
594
+ is still theirs to make.
595
+
596
+ Then stop. The next step is `{{command_prefix}}make`.
597
+
598
+ ## make
599
+
600
+ **High-fidelity.** Build the actual page or feature from its confirmed spec and its
601
+ approved mockup — the real tokens, components and copy, working.
602
+
603
+ **After a critique,** `make` is how its findings get fixed: read the report, fix
604
+ each finding, run the finish again, and then `critique` runs again. See step 5 of
605
+ `critique`.
606
+
607
+ ### Refuse rather than guess
608
+
609
+ - No `.jig/specs/<surface>.spec.md` → say so and run `{{command_prefix}}spec`
610
+ first. Do not write a quick spec of your own to satisfy the check; a spec you
611
+ wrote and accepted in the same breath is the thing the confirmation exists to
612
+ prevent.
613
+ - `confirmed: false` → the spec exists but nobody has agreed to it. Ask for
614
+ confirmation and stop.
615
+ - `mockup: pending` → the feature has not been looked at. Run
616
+ `{{command_prefix}}mockup`, or ask the user whether to skip it. Do not decide to
617
+ skip it yourself.
618
+
619
+ ### Build
620
+
621
+ **Build from the spec and the mockup, never by copying the mockup.** The spec says
622
+ what is being built; the approved mockup shows its structure at each size — what
623
+ comes first, what is grouped with what, what stacks on the phone. Match that
624
+ structure at high fidelity.
625
+
626
+ What you do not do is paste the mockup's HTML into the product, or take the code a
627
+ Figma or Stitch frame will generate for you. The drawing has fixed frame widths, grey
628
+ boxes and no tokens; pasted in, it carries none of the rules or tokens the product is
629
+ built on. Read it; do not copy it.
630
+
631
+ If the spec and the approved mockup disagree, the spec was not updated during review.
632
+ Stop and ask the user which is right, fix the spec, and continue — do not choose
633
+ between them yourself.
634
+
635
+ If `mockup: skipped`, build from the spec alone.
636
+
637
+ Build V1 only. Nothing in `later:`, however little extra it looks.
638
+
639
+ Build the **phone** composition first, then add what `tablet` and `desktop`
640
+ change. Write base styles for the phone and add width as it is available — not
641
+ the desktop layout with overrides that take it apart again. A layout built wide
642
+ and subtracted from ends up correct at exactly the widths somebody checked.
643
+
644
+ Then the ordinary rules apply: tokens by semantic name, the relevant
645
+ `03-patterns.md` section for each component, `05-copy.md` for every string.
646
+
647
+ ### Record what you changed
648
+
649
+ Building reveals that a spec was wrong somewhere — that is normal and expected.
650
+ **Silence is the failure, not the change.** Append to `deviations:`:
651
+
652
+ ```yaml
653
+ deviations:
654
+ - field: sizes.phone.form.steps
655
+ from: 1
656
+ to: 2
657
+ why: "password reset needs an email round-trip; one screen cannot hold both"
658
+ ```
659
+
660
+ A spec that quietly stops describing the page is worse than no spec: a later
661
+ review then validates the page against fiction and reports it clean.
662
+
663
+ ### Finish
664
+
665
+ You are not done until all three of these hold. They are steps, not advice: in a
666
+ live run one build skipped `check` and shipped 53 errors, another offered 36 as
667
+ "non-blocking", and a third drew no phone navigation where its approved mockup had a
668
+ menu and reported "matched exactly".
669
+
670
+ **1. `check` passes.** Run `{{scripts_path}} check --all` and paste its
671
+ `JIG_CHECK:` line into your report, as printed. `mechanical=fail:<n>` means not done:
672
+ fix every error and run it again. No mechanical finding is non-blocking, and you do
673
+ not get to decide one is. Warnings are listed and either fixed or explained one by
674
+ one. A report with no `JIG_CHECK:` line, or one you typed yourself, is not a finished
675
+ build.
676
+
677
+ **2. The build matches the approved mockup at each size.** Skip only when
678
+ `mockup: skipped`. Render the page at 360px, 768px and 1280px — with a browser, not in
679
+ your head — and set each render beside the mockup at the same width. Go region by
680
+ region, in the mockup's order:
681
+
682
+ | Size | Mockup region | In the build? | Same place and order? |
683
+ | --- | --- | --- | --- |
684
+ | phone | header: name + Menu button | yes | yes |
685
+ | phone | plan cards, stacked | yes | no — Pro first |
686
+
687
+ Every **no** is either fixed now or appended to `deviations:` with a `why`. Never
688
+ write "matches" without the table: the claim is what the table is for. A control the
689
+ mockup draws — a Menu button, a toggle — is in the build only if it works: tap it.
690
+
691
+ **3. Each size's spec fields are accounted for.** Say which fields the built page
692
+ satisfies and which it does not — **at each size**. A page that satisfies `desktop`
693
+ and was never looked at on a phone has satisfied a third of its spec. Then run the
694
+ self-check at the end of `{{rules_path}}/00-anti-patterns.md`.
695
+
696
+ Report the `JIG_CHECK:` line, the table, and any new deviations. Then suggest
697
+ `{{command_prefix}}critique`.
698
+
699
+ ## critique
700
+
701
+ Scrutinise what was built against the rules and its spec. This is **not** a second `check`, and the difference is
702
+ not effort — it is evidence.
703
+
704
+ `check` reads files. `critique` looks at the page **and** reads the source
705
+ against the rules `check` has no detector for. A hierarchy failure is in no
706
+ file: it is in what the files produce together, so no amount of careful reading
707
+ finds it. That is what the render arm is for.
708
+
709
+ Every judgment rule says which pass owns it, in `pass:`. The corpus divides
710
+ three ways, and each part needs a different instrument:
711
+
712
+ | | rules | instrument |
713
+ |---|---|---|
714
+ | mechanical, has a `detector` | a small minority | `check`, deterministic |
715
+ | judgment, `pass: screen` | the render | arm A, looking at the page |
716
+ | judgment, `pass: code` | **the largest group** | arm C, reading the source |
717
+
718
+ The third one is the reason this command exists in its current form. `check`
719
+ judges only what a detector can decide, which is a fraction of the corpus. Every
720
+ other `pass: code` rule is judgment, and until something reads the source against
721
+ it, it is enforced by nothing but the building agent having remembered it. An
722
+ agent under load does not remember 60-odd rules; it consults the handful it
723
+ thought to search for, and ships the rest unchecked.
724
+
725
+ No two arms judge the same rule, which is why they can never hand a user
726
+ competing verdicts on one id.
727
+
728
+ ### Refuse rather than guess
729
+
730
+ - No `DECISIONS.md` with substance → run `{{command_prefix}}decide` first. `critique`
731
+ judges the page against the project's decisions as well as the rules, and has
732
+ nothing to judge that half against without them.
733
+ - The user named nothing → ask which page, feature or functionality.
734
+ - There is no `.jig/specs/<surface>.spec.md` → say so and continue **rules-only**,
735
+ reporting `spec=missing`. Do not invent the spec the page should have had; a
736
+ spec written after the page, from the page, agrees with it by construction.
737
+
738
+ ### 1. Three assessments that cannot see each other
739
+
740
+ Run all three. **No arm may see another's output**, and none may see the
741
+ conversation that built the page. Running them in one head anchors them to each
742
+ other; do not shortcut it for cost, speed or context size.
743
+
744
+ - **A — the reader.** Delegate to a subagent. Give it the URL or file, the spec,
745
+ the approved mockup (from `mockup_at:`), `DECISIONS.md`, and the corpus. It judges the `pass: screen` rules against the render, **and the `P-` patterns the spec's regions use** — a navigation region means `P-14`. Those patterns are not in `rules.index.json`, so walking the index alone never reaches them. It writes its verdicts to `.jig/critique/<surface>/screen.json` (step 1c). Do
746
+ **not** give it the build conversation. If it needs that conversation, it is
747
+ not reviewing the page, it is agreeing with itself.
748
+ - **B — the machine.** Run `check`. Deterministic, same input same output. It
749
+ judges the rules that carry a `detector`, and nothing else — its own
750
+ attestation says `judgment=not-run`, and that is a true statement about the
751
+ rest of the corpus, not a formality.
752
+ - **C — the source reader.** A second subagent, separate from A. Give it the
753
+ source files, the spec, and the corpus, and no render. It judges every
754
+ `pass: code` judgment rule — the group `check` cannot reach — and writes its
755
+ verdicts to `.jig/critique/<surface>/code.json` (step 1c).
756
+
757
+ **If you cannot delegate an arm, it is skipped. You do not run it yourself.** No
758
+ subagent available, a subagent that cannot read the project's files — either way
759
+ the arm reports `skipped` with the reason, in words, in the report. A live run had
760
+ an agent build a page and then run the reader arm on it in its own head; every
761
+ verdict it wrote cited what the spec *intended* rather than what the page *did*,
762
+ and its one finding asked the builder to confirm it. A skipped arm is an honest
763
+ gap. A self-run arm is a false review that looks like a real one.
764
+
765
+ Each reader is a subagent for the same reason a writer does not proofread their
766
+ own copy. An agent that built the page knows what it meant, and what it meant is
767
+ invisible in the artifact — which is the one thing a reviewer must not have.
768
+
769
+ A and C are kept apart for a narrower reason: an arm that has seen the render
770
+ stops reading the source and starts confirming the picture. They answer to
771
+ different evidence and must not share it.
772
+
773
+ ### 1a. Arms A and C walk the index. They do not search it.
774
+
775
+ **This is not a style preference. It is the difference between covering the
776
+ corpus and sampling it.**
777
+
778
+ Load `rules.index.json`, filter to the pass the arm owns, and return a verdict
779
+ for **every id in that list**: `ok`, `finding`, or `n/a` with one line of
780
+ reason. A rule that does not apply is still answered — `n/a, this page has no
781
+ form` is a verdict.
782
+
783
+ Searching the corpus is how a *builder* works, and it is the right method there:
784
+ it finds the rules you can already name. It cannot find the rule for the mistake
785
+ you do not know you are making, which is the only kind of rule worth writing
786
+ down. A pass driven by search returns what the agent thought to look for, and
787
+ reports nothing about everything else — indistinguishable, in the output, from a
788
+ clean page.
789
+
790
+ The builder cannot do this. It is holding a brief, a spec and a half-written
791
+ page, and it will drop most of the corpus under that load. An arm with one job
792
+ can walk all of it, which is the whole reason the work is split.
793
+
794
+ **Never narrow the list by guessing relevance before reading.** Filter by `pass`,
795
+ and by mode where a rule is mode-gated. Nothing else. An arm that decides in
796
+ advance which rules are worth checking has reintroduced search with extra
797
+ steps.
798
+
799
+ ### 1b. What makes a verdict real
800
+
801
+ Every one of these was produced by a live run of this command, by arms that had
802
+ been told to walk the index and believed they were doing it.
803
+
804
+ **An `n/a` must name what the rule is about.** Not "does not apply" — *what* does
805
+ not apply. `n/a, D-27 is about optical alignment and this page has no icons` is a
806
+ verdict. `n/a, rule context unavailable` is an absence wearing a verdict's shape,
807
+ and the difference is mechanically checkable: run
808
+ `{{command_prefix}}explain <id>` and the rule either resolves or it does not.
809
+
810
+ One run returned `C-68 — rule not found in accessible corpus` and
811
+ `D-27 — rule context unavailable`. Both resolve instantly. A third described
812
+ `D-25` as being about label spacing; it is about proximity hierarchy. None of
813
+ the three had been read. `C-68` is *non-interactive elements styled like
814
+ interactive ones* — adjacent to the only real finding that arm had raised.
815
+
816
+ **An id you cannot resolve is a hard stop.** If `explain` fails on an id that is
817
+ in the index, that is a corpus or install defect. Say so, loudly, and stop.
818
+ Recording it as a rule that happens not to apply buries a broken install inside a
819
+ clean-looking report.
820
+
821
+ **Never treat the count as a target.** The instruction is *return a verdict for
822
+ every id*; `<n>` is the consequence. An arm asked to reach a number will reach
823
+ it — that is how the three fabrications above were produced, in one turn, after a
824
+ reviewer asked for a short pass to be completed. **A pass that comes back short
825
+ is re-run, not topped up.**
826
+
827
+ **An arm that contradicts itself is re-run, not merged.** A live run's reader
828
+ arm said, in one report, that a toggle hard-codes `60px`/`34px` instead of tokens
829
+ (`A-07`) and that nothing on the page is hard-coded (`H-47`). Two opposite verdicts
830
+ on the same file are not a disagreement to report; they are an instrument that
831
+ cannot be trusted, and every other verdict it returned is suspect with them.
832
+
833
+ ### 1c. Verdicts go in files, and the CLI counts them
834
+
835
+ Each reader arm writes its own file. The arm reports back only that it wrote it.
836
+
837
+ ```json
838
+ {
839
+ "rendered": true,
840
+ "artefacts": [".jig/critique/pricing/360.png", ".jig/critique/pricing/768.png", ".jig/critique/pricing/1280.png"],
841
+ "verdicts": [
842
+ { "id": "D-115", "verdict": "ok", "reason": "no sideways scroll at 360, 768 or 1280" },
843
+ { "id": "P-14", "verdict": "finding", "reason": "menu button at 360 does not open; aria-expanded never set" },
844
+ { "id": "A-60", "verdict": "n/a", "reason": "A-60 is about competing icons; this page has none" }
845
+ ]
846
+ }
847
+ ```
848
+
849
+ `screen.json` carries `rendered` and `artefacts`; `code.json` needs only `verdicts`.
850
+
851
+ **A rendered review is measured, not only described.** At 360, 768 and 1280px, run the
852
+ probe in the browser and save what it returns beside the verdicts:
853
+
854
+ ```
855
+ {{scripts_path}} probe > .jig/probe.js
856
+ # at each width: set the viewport, open the page, evaluate .jig/probe.js, save the
857
+ # JSON it returns as .jig/critique/<surface>/probe-<width>.json
858
+ ```
859
+
860
+ The probe opens the phone menu, presses Escape, measures sideways scroll, and reads
861
+ whether the styles and tokens actually applied. `verdicts` refuses `rendered: true`
862
+ without a probe at each width, and refuses an `ok` the probe contradicts — a `P-14`
863
+ "ok" on a menu that did not open, a `D-115` "ok" on a page that scrolls sideways. It
864
+ also fails a page that renders in the browser's default font, or shows `${` to
865
+ readers, whatever the verdicts say. Never write a probe file yourself: when it
866
+ disagrees with you, it is the review that is wrong.
867
+ Then run:
868
+
869
+ ```
870
+ {{scripts_path}} verdicts <surface>
871
+ ```
872
+
873
+ **It decides whether the review is complete, not you.** It fails on an id that is not
874
+ in the corpus, names every id in the arm's pass that has no verdict, rejects a rule
875
+ filed by the wrong arm or judged twice, rejects an `n/a` whose reason says the rule
876
+ was not read, requires `P-14` when the spec has navigation, and refuses
877
+ `rendered: true` without artefacts that exist. It ends with a `JIG_VERDICTS:` line.
878
+
879
+ When it fails, **re-run the arm it names**. Do not edit the file to make it pass, and
880
+ do not report an incomplete arm as having run. A live Haiku run of this command
881
+ before it existed produced invented rule ids, `screen=ran:1` filed as a review,
882
+ `code=ran:97` against 67, `ran:100` against both, and arms that reported to the
883
+ wrong agent entirely. Every one of those was forbidden here in prose. Prose did not
884
+ hold; a check does.
885
+
886
+ The merge reads the two files — not a summary an arm sent back. An arm's message can
887
+ go astray; its file cannot.
888
+
889
+ ### 2. Look at the page
890
+
891
+ If a browser is available this is **required**, not optional. Use whatever the
892
+ harness provides — a browse skill, Playwright, a headless browser. Reading the
893
+ source and imagining the page is not a render.
894
+
895
+ - **Operate the page, don't only look at it.** At 360px, open the navigation
896
+ menu: the links must appear, `aria-expanded` must change, and the control's label
897
+ or icon must show that it now closes. Use the billing control, the FAQ, anything
898
+ that changes state, and check it changes. A menu button that renders is not a menu
899
+ that works: in a live run four of six pages shipped a phone menu that did not open,
900
+ or no menu at all, and every review that only looked passed them.
901
+ - **Render at every size the spec names** — phone at 360px, tablet at 768px,
902
+ desktop at 1280px — and judge each size's fields against its own render. A
903
+ `same-as:` claim is checked like any other field: render that size and see
904
+ whether the composition really does hold.
905
+ - At every size, measure `document.documentElement.scrollWidth` against
906
+ `clientWidth`. Greater means the page scrolls sideways at that width, which is a
907
+ finding however good the rest of it looks.
908
+ - Compare the sizes. Did the phone change the **composition**, or only the
909
+ **dimensions**? Reordering, stacking, collapsing, a different control — that is
910
+ composition. The same layout with smaller numbers is not.
911
+ - Check the page is styled at all. A stylesheet that 404s renders a page that
912
+ passes every file-based check ever written.
913
+ - Run the squint test from `L-01` against the render, not the analogue: render
914
+ each size once more with `filter: grayscale(1)` on the root, and check that the
915
+ primary action, the headings and the groups still read in order without colour.
916
+ Keep the grayscale screenshots with the others; they are artefacts too. The
917
+ analogue exists for when there is no browser.
918
+
919
+ If no browser is available, say so in one line and use `L-01`'s analogue. Never
920
+ skip it silently, and never report a render you did not do.
921
+
922
+ ### 3. Compare the page to its spec and its mockup
923
+
924
+ **The spec**, field by field: for each size, say whether the page rendered at that
925
+ size satisfies each frontmatter field.
926
+
927
+ **The mockup**, size by size: set each rendered size beside the approved drawing
928
+ at the same width and compare the structure — the regions and their order, what is
929
+ grouped with what, what stacks or moves on the phone. Not the colour, type or
930
+ polish: the mockup is grayscale and low-fidelity on purpose, so its appearance is
931
+ not a target.
932
+
933
+ A difference from either is a finding **unless** `deviations:` already explains it
934
+ — that is what the recorded deviation is for. A difference explained by neither is
935
+ the more serious finding, because it means the spec has quietly stopped describing
936
+ the page. If `mockup: skipped`, compare to the spec alone and say so.
937
+
938
+ ### 4. Report
939
+
940
+ Merge both assessments into one report, keyed by rule id, ordered by severity.
941
+ Say explicitly where the two agree, what only the machine found, what only the
942
+ reader found, and which machine findings are false positives. **The disagreement
943
+ is information, not noise to smooth over.**
944
+
945
+ Say what is right as well as what is wrong — two or three things, and why they
946
+ work. A review that only lists faults tells the reader nothing about what to
947
+ preserve while fixing them.
948
+
949
+ Then the attestation:
950
+
951
+ ```text
952
+ JIG_CRITIQUE: version=<version> mode=<mode> surface=<name> spec=<confirmed|missing|unconfirmed> mockup=<approved|skipped|missing> rendered=<yes|no> screen=<ran|skipped>:<n> code=<ran|skipped>:<n> mechanical=<pass|fail|skipped>:<n> warnings=<n>
953
+ ```
954
+
955
+ **Take `screen=`, `code=` and `rendered=` from `jig verdicts`, never from your own
956
+ count.** If it reported an arm incomplete, that arm is `skipped` here with the reason,
957
+ until it is re-run and passes.
958
+
959
+ Three counts because there are three arms, and each can fail independently.
960
+ `mechanical=` is `check`'s result and uses `check`'s own field name deliberately,
961
+ so the two records line up. `screen=` and `code=` are the two reader arms, and
962
+ their `<n>` is **the number of rules the arm returned a verdict for** — not the
963
+ number of findings. An arm that walked the index reports the size of the list it
964
+ walked; an arm reporting far fewer did not walk it, and the number is where that
965
+ shows.
966
+
967
+ `code=skipped:0` is the state this command was rebuilt to make visible. It used
968
+ to be unsayable: the largest group of rules in the corpus went unjudged and the
969
+ report looked complete.
970
+
971
+ `rendered=no` is not a detail. Without it a clean report reads as "this page is
972
+ good" when it can only mean "the files are good" — and a page whose stylesheet
973
+ never loaded satisfies every file-based check there is.
974
+
975
+ `rendered=yes` means **every size the spec names** was rendered. If only some
976
+ were, the report names which, and says what was not seen.
977
+
978
+ **`rendered=yes` requires an artefact.** A screenshot, a rendered dump, something
979
+ that exists because a browser ran. Reading the source and reasoning about how it
980
+ would look is the `L-01` analogue, and the analogue is `rendered=no`. A live run
981
+ of this command emitted `rendered=yes` on a machine with no browser installed and
982
+ no image written anywhere — turning an honest limitation into a false assurance,
983
+ on the one field a reader leans on hardest. If you did not see the page, say you
984
+ did not see the page.
985
+
986
+ ### 5. Then the cycle goes round again
987
+
988
+ A critique with findings is not the end of the feature. It is the middle of it.
989
+ Designing every case in advance does not work; fixing a page you can use does. So:
990
+
991
+ 1. Tell the user the findings, and that the next step is `{{command_prefix}}make`
992
+ to fix them — not the next feature.
993
+ 2. `make` fixes them, runs its finish again, and records any spec change in
994
+ `deviations:`.
995
+ 3. Run `{{command_prefix}}critique` again on the same surface.
996
+ 4. Repeat until the report has no errors or warnings, or the user has accepted each
997
+ one that is left **by id, in their own words** ("leave E-65 for now"). Your
998
+ judgment that a finding is minor is not acceptance.
999
+
1000
+ Only then suggest `{{command_prefix}}spec` for the next feature. Starting the next one
1001
+ with findings open means the second feature is built on top of the first one's
1002
+ problems, and nobody goes back for them.
1003
+
1004
+ ## probe
1005
+
1006
+ Run `{{scripts_path}} probe`. It prints one JavaScript expression — the render probe —
1007
+ for the critique's screen arm to evaluate in a browser at each width. See step 1c of
1008
+ `critique`. It changes nothing on its own.
1009
+
1010
+ ## gate
1011
+
1012
+ Not a command to run. In Claude Code, `install` adds a Stop hook that runs
1013
+ `{{scripts_path}} gate` whenever you try to finish. If files you changed fail `check`,
1014
+ or a critique's verdict files fail `verdicts`, it refuses to let you stop and hands
1015
+ you the failures as the reason. **Fix what it names.** Do not edit verdict files to
1016
+ satisfy it, and do not report the work as done while it is blocking. After three
1017
+ refusals in one session it lets you stop; then tell the user plainly that the work is
1018
+ not finished and what is still failing.
1019
+
1020
+ It exists because every "run check" and "run verdicts" step in these procedures was
1021
+ skipped in a live run, and pages that did not render their styles were reported
1022
+ clean.
1023
+
1024
+ ## verdicts
1025
+
1026
+ Run `{{scripts_path}} verdicts <surface>`, passing the argument through, and report its
1027
+ output unchanged — every `✗` line and the closing `JIG_VERDICTS:` line.
1028
+
1029
+ It exists so that `critique`'s counts are computed rather than written. Read its
1030
+ errors as instructions to re-run the arm they name. It never fixes a verdict file,
1031
+ and neither do you: a file edited until it passes records nothing about the page.
1032
+
1033
+ ## decide
1034
+
1035
+ Write `DECISIONS.md` beside the token layer — `jig/DECISIONS.md` by default, or
1036
+ wherever `brand` in `{{config_file}}` puts the token files.
1037
+
1038
+ `spec`, `mockup`, `make` and `critique` are blocked until this exists: `spec` and
1039
+ `critique` refuse to start without it, and `mockup` and `make` need a confirmed spec.
1040
+ Nothing else is blocked — `install`, `init`, `check` and `explain` run without it.
1041
+ `init` comes first, because this file sits beside the token files `init` writes.
1042
+
1043
+ ### What belongs in it, and what does not
1044
+
1045
+ The token layer already holds **values**, and `{{config_file}}` already holds
1046
+ **mode**. Neither can hold a **reason**. This file is the reasons.
1047
+
1048
+ - ✅ *"Our accent arrives as a whole field or not at all — a full-bleed section,
1049
+ a solid button, a filled state. Never a 2px underline, never an icon tint.
1050
+ Its authority comes from arriving in quantity, rarely."*
1051
+ - ❌ *"The accent is `#FF6347`."* — that is a token. Restating it here creates
1052
+ two places to change it, and they will disagree.
1053
+ - ❌ *"Use 4.5:1 contrast."* — that is `C-19`, and it is true of every project.
1054
+ A universal rule copied into a project file is a rule that can rot locally.
1055
+
1056
+ The test: **could this be derived from the tokens, the mode, or a numbered rule?**
1057
+ If yes, it is already written down somewhere better. If no, it is a decision, and
1058
+ this is where it lives.
1059
+
1060
+ **`decide` is product-wide. `spec` is per screen.** Ask what holds on every
1061
+ screen. What one screen contains, who arrives at it, what its FAQ asks, where
1062
+ something sits on it — those are `spec`'s questions. Asking them here makes the
1063
+ user answer them twice, and a live run did exactly that: one agent gathered a
1064
+ page's FAQ and arrival paths under `decide`, the other under `spec`.
1065
+
1066
+ One kind of placement **is** a decision: something that must sit in the same
1067
+ place on every screen, such as where the menu button lives. That is taste, not
1068
+ something a rule can settle, and it belongs here once. Write it as where the thing
1069
+ sits **when it appears** — "the Menu button, wherever it appears, sits top right" —
1070
+ never as "a menu button on every page". Whether a menu button exists at a width is
1071
+ `P-14`'s, decided per size in `spec`: at widths where the links fit there is none. A
1072
+ live run recorded "Menu button: top right, same on every page", and all four specs
1073
+ read it as a menu button at desktop width too. Placement on a single
1074
+ screen is not — and do not ask the owner where an element should go on a screen
1075
+ at all. If the design system has a position, it is in the rules; if it has none,
1076
+ `spec` settles it.
1077
+
1078
+ ### How to write it
1079
+
1080
+ **Interview, do not infer.** Two or three questions per round, then wait. What
1081
+ you are looking for is the places where this project has already chosen
1082
+ something, or will have to:
1083
+
1084
+ - Round 1 — what is this product, and who is it for in a sentence the team would
1085
+ recognise? What should it never look like? Name real products, not adjectives.
1086
+ - Round 2 — for each thing the tokens already set, is there a rule about *how* it
1087
+ is used? Where does the brand colour go, and where is it forbidden? What earns
1088
+ emphasis?
1089
+ - Round 3 — what has the team already argued about, or reversed? A decision with
1090
+ a history is the one most worth recording, because it is the one most likely to
1091
+ be re-made wrongly.
1092
+
1093
+ **All three rounds run.** Do not write the file after round 2 because the answers
1094
+ look complete. Round 3 asks for the one kind of decision a user does not volunteer
1095
+ unprompted — the ones with a history — and a live run skipped it for exactly that
1096
+ reason. In the arm that did ask, round 3 is where the unresolved items surfaced.
1097
+
1098
+ **Round 3 also asks, by name, what is still undecided.** "Is there anything the team
1099
+ has not settled yet — something still argued about, or waiting on someone?" Ask it
1100
+ even when nothing so far suggests there is. Whatever the answer names goes in the
1101
+ **Unresolved** section below, not into a decision the user did not make. Four of six
1102
+ files in a live run dropped the one open item the owner had named.
1103
+
1104
+ **Write each decision as a named rule with its reason attached**, in the project's
1105
+ own vocabulary:
1106
+
1107
+ ```markdown
1108
+ ### The Stamp Rule
1109
+
1110
+ The accent appears as a whole field or not at all. A full-bleed section, a solid
1111
+ button, a filled active state — never a 2px underline, never a small icon tint,
1112
+ never a gradient stop. Its authority comes from arriving in quantity, rarely.
1113
+
1114
+ **Why:** thinly spread, it reads as decoration and stops meaning anything. The
1115
+ audit test: if a coloured element is smaller than a section band, it is probably
1116
+ wrong.
1117
+ ```
1118
+
1119
+ A name gives the team something to cite in review. A reason lets a future agent
1120
+ tell when the rule does not apply, which a bare instruction never can.
1121
+
1122
+ **The reason is the owner's, or it is not written.** Write the `Why:` the user gave,
1123
+ in their words or close to them. If they gave none, ask for it. If they still give
1124
+ none, write `**Why:** not given` — never a reason you supplied. A live run invented
1125
+ reasons for two real reversals; an invented reason is worse than none, because the
1126
+ next agent weighs it as the team's and applies the rule where the team never would.
1127
+
1128
+ **Record what is still open in its own section, at the end:**
1129
+
1130
+ ```markdown
1131
+ ## Unresolved
1132
+
1133
+ - **Annual discount on the pricing page.** Undecided between showing it on the
1134
+ toggle and only at checkout; waiting on finance. Until settled, specs do not
1135
+ commit to either — they ask.
1136
+ ```
1137
+
1138
+ Each item says what is undecided, between what, and what it is waiting on. This
1139
+ section is not a placeholder and does not fail the gate: an honest "not decided yet"
1140
+ is information. `spec` reads it, and asks the user instead of choosing when a screen
1141
+ touches an unresolved item. If round 3 surfaced nothing, write `None named by the
1142
+ owner.` under the heading, so a reader can tell it was asked.
1143
+
1144
+ ### Finish
1145
+
1146
+ **Before you show it, account for every answer.** List each thing the user said,
1147
+ round by round, and the section of the file it went to — a named decision, its
1148
+ `Why:`, or **Unresolved**. An answer that went nowhere is either added or shown to
1149
+ the user as left out on purpose, with the reason. Then go the other way: every
1150
+ `Why:` in the file must trace to something the user said; delete any that does not,
1151
+ and ask. Include this mapping in your message, below the file.
1152
+
1153
+ Show it and ask the user to confirm it is right before you stop. Leave no
1154
+ `[TODO]` markers — the gate treats them as an unwritten file, correctly, because
1155
+ a placeholder decision is not a decision.