task-pipeline-skill 1.88.1 → 1.89.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +64 -0
- package/README.md +1 -1
- package/SKILL-CARD.md +1 -1
- package/package.json +3 -3
- package/plugins/task-pipeline/.claude-plugin/plugin.json +1 -1
- package/plugins/task-pipeline/agents/verifier-product.md +3 -1
- package/plugins/task-pipeline/agents/verifier-visual.md +115 -0
- package/plugins/task-pipeline/agents/verifier.md +2 -1
- package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +4 -4
- package/plugins/task-pipeline/skills/task-pipeline/graph.schema.json +22 -1
- package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +4 -4
- package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +12 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/audit.md +14 -8
- package/plugins/task-pipeline/skills/task-pipeline/references/browser.md +97 -3
- package/plugins/task-pipeline/skills/task-pipeline/references/certification.md +46 -5
- package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +28 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/conventions.md +3 -3
- package/plugins/task-pipeline/skills/task-pipeline/references/doctrine-map.md +1 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/grill.md +15 -9
- package/plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md +23 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/portability.md +1 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/spec.md +27 -3
- package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +75 -13
- package/plugins/task-pipeline/skills/task-pipeline/references/work-graph.md +2 -2
- package/plugins/task-pipeline/skills/task-pipeline/scripts/graph.py +45 -16
- package/plugins/task-pipeline/skills/task-pipeline/scripts/stage_checkpoint.py +21 -1
- package/plugins/task-pipeline/skills/task-pipeline/scripts/visual_gate.py +589 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/brief.md +5 -2
- package/plugins/task-pipeline/skills/task-pipeline/templates/browser-claims.json +223 -1
- package/plugins/task-pipeline/skills/task-pipeline/templates/run.md +10 -0
package/CHANGELOG.md
CHANGED
|
@@ -1,3 +1,67 @@
|
|
|
1
|
+
## v1.89.1 — v1.89.0's payload, released with its no-stamp declaration
|
|
2
|
+
|
|
3
|
+
`v1.89.0` was refused by its own release check: the release carried no run stamp and was not
|
|
4
|
+
named in `## Releases that carry no stamp`. Nothing was published under it. This release is the
|
|
5
|
+
same payload with the declaration in `docs/evidence/retro.md`, the tenth such instance. What
|
|
6
|
+
v1.89.0 ships is described in its section below.
|
|
7
|
+
|
|
8
|
+
Guards: 430 → **430** — a declaration and a version bump; no check changes.
|
|
9
|
+
|
|
10
|
+
## v1.89.0 — the visual layer is checked by its trace, and a fourth reading opens the picture
|
|
11
|
+
|
|
12
|
+
Until now the stage-3 VISUAL track passed on the fact that it ran, the look at stages 5–6
|
|
13
|
+
read the accessibility tree and never the pixels, and none of the three verifiers opened a
|
|
14
|
+
picture. A landing could pass every gate while nobody compared what shipped with what was
|
|
15
|
+
designed. This release adds the visual half, gated by a class the brief now records
|
|
16
|
+
(`DEC-0006`).
|
|
17
|
+
|
|
18
|
+
Guards: 430 → **430** — the new checks are suites of their own and the validator is unchanged.
|
|
19
|
+
|
|
20
|
+
- **Stage 0 records `surface_class: flagship | product | internal | ad`** for every
|
|
21
|
+
user-facing task (`references/stages.md` → *The surface class*, the grill's UI branch,
|
|
22
|
+
`templates/brief.md`). The class selects the gate profile for stages 3, 6 and 10.
|
|
23
|
+
- **Stage 3 reads the director record's fields, not the fact the track ran.**
|
|
24
|
+
`scripts/visual_gate.py record <file> --class <c>` checks the headings each class owes
|
|
25
|
+
(the full set on `flagship`; Brief, Mode, References, Markers, Open on `product`) and runs
|
|
26
|
+
sheleg-design's `--check-record` where it is installed and new enough. Where it is not,
|
|
27
|
+
the validator reads NOT_RUN beside the verdict, never PASS. `Mode: declined` with a reason
|
|
28
|
+
is the refusal, and it passes (`references/spec.md`).
|
|
29
|
+
- **Figma: one file per surface (App, Web, ASO), not one for every frame.**
|
|
30
|
+
`visual_gate.py filekeys` checks every `screens.md` frame link's key against the recorded
|
|
31
|
+
set (`grill.md`, `stages.md`, `audit.md`'s `→F` seam).
|
|
32
|
+
- **The visual half of the look** (`references/browser.md` → *The visual half*). Checks
|
|
33
|
+
run cheapest first: the project linter (`visual_gate.py lint`, NOT_RUN exit 3 where
|
|
34
|
+
absent), regression against the baseline, then the state × axes matrix (pairwise plus
|
|
35
|
+
the mandatory pairs dark × large text and RTL × narrow). Each frame carries its capture
|
|
36
|
+
record and a diff against its Figma frame or baseline. A judge reads the rubric. J items
|
|
37
|
+
stay NOT_ASSESSED until a labelled set calibrates the judge, and no judge overrides a G
|
|
38
|
+
FAIL. It is a gate on `flagship`, `product` and `ad`, and recommended on `internal`.
|
|
39
|
+
Re-render budget: one, two at most, then `unresolved` to the person. Only external,
|
|
40
|
+
specific feedback starts a round (`loop-guard.md` → *The re-render loop*).
|
|
41
|
+
- **The contact sheet is `templates/browser-claims.json`, extended rather than replaced**
|
|
42
|
+
— row fields `axes`, `capture`, `figma_frame`, `baseline`, `diff`, `rubric[]`; file
|
|
43
|
+
fields `surface`, `revision`, `review_rounds`, `approved_by`, `approved_at`.
|
|
44
|
+
`visual_gate.py sheet` validates it and exits 0 · 1 · 2 · 3. The browser-claims rules now
|
|
45
|
+
live in that shipped script, so a host can run them. `test/browser_claims_test.py`
|
|
46
|
+
imports them under the old names.
|
|
47
|
+
- **`agents/verifier-visual.md` — the fourth blind reading.** `graph.py certify` requires
|
|
48
|
+
a `visual` report on a node whose new `surface_class` is `flagship` or `product`, and
|
|
49
|
+
accepts one on any other node. `graph.schema.json` gains the field and the tier
|
|
50
|
+
(`references/certification.md` → *The fourth reading*).
|
|
51
|
+
- **Acceptance walks visual intent ↔ the approved contact sheet** (`audit.md`'s `V→R`
|
|
52
|
+
seam). The run ledger gains a `review:` line. `stage_checkpoint.py` carries it into the
|
|
53
|
+
checkpoint of the stage it names, so the passes a surface took are measured.
|
|
54
|
+
- **Visual lanes in `companion-skills.md`.** These are tools, never entry points:
|
|
55
|
+
`break-ui`, `review-animations` / `improve-animations`, `mobile-native`, `animate-expo`,
|
|
56
|
+
`webapp-testing` / `chrome-devtools`, `accessibility-review` / `a11y-debugging`, and the
|
|
57
|
+
platform audits.
|
|
58
|
+
- **Tests:** `test/visual_gate_test.py` (24 cases); `test/browser_claims_test.py` (35
|
|
59
|
+
cases, 27 of them new contact-sheet plants); six `graph_test.py` cases for the visual tier and
|
|
60
|
+
the class; one `stage_checkpoint_test.py` case; and a `certify_mutations.py` mutation for
|
|
61
|
+
the visual requirement. A hand mutation pass over `visual_gate.py` disabled each of 34
|
|
62
|
+
rules in turn, and a fixture failed for every one. CI now also runs the three suites
|
|
63
|
+
`npm test` already ran.
|
|
64
|
+
|
|
1
65
|
## v1.88.1 — v1.88.0's payload, released with its no-stamp declaration
|
|
2
66
|
|
|
3
67
|
The same payload as `v1.88.0`: the stage-boundary checkpoint writer, `scripts/stage_checkpoint.py`
|
package/README.md
CHANGED
|
@@ -182,7 +182,7 @@ until it is installed.
|
|
|
182
182
|
| 3 Spec | [`spec.md`](plugins/task-pipeline/skills/task-pipeline/references/spec.md) — UX-track order, locked contracts, global constraints, self-review |
|
|
183
183
|
| 4 Plan | [`planning.md`](plugins/task-pipeline/skills/task-pipeline/references/planning.md) — zero-context tasks, parallel groups, no placeholders |
|
|
184
184
|
| the queue | [`work-graph.md`](plugins/task-pipeline/skills/task-pipeline/references/work-graph.md) — a script walks the graph so the model never reads it: 400 nodes and 4 print the same 27-byte frontier |
|
|
185
|
-
| closing a node | [`certification.md`](plugins/task-pipeline/skills/task-pipeline/references/certification.md) — three blind readings at escalating visibility (the changed code, what reaches it, the product around it), all three required to pass; a failing round records itself and the node stays open |
|
|
185
|
+
| closing a node | [`certification.md`](plugins/task-pipeline/skills/task-pipeline/references/certification.md) — three blind readings at escalating visibility (the changed code, what reaches it, the product around it), all three required to pass — and a fourth, `verifier-visual`, on a flagship or product surface, which reads the contact sheet against the director record; a failing round records itself and the node stays open |
|
|
186
186
|
| 5 Build | [`build.md`](plugins/task-pipeline/skills/task-pipeline/references/build.md) + [`review.md`](plugins/task-pipeline/skills/task-pipeline/references/review.md) — isolation, ledger, subagent loop, review rubric, fix loop |
|
|
187
187
|
| 5–6 TDD | [`tdd.md`](plugins/task-pipeline/skills/task-pipeline/references/tdd.md) — the iron law, red/green/refactor, the suite gate |
|
|
188
188
|
| 5, 6, 8 The browser | [`browser.md`](plugins/task-pipeline/skills/task-pipeline/references/browser.md) — the ref model both channels share, the four commands the look is made of, sessions, and the three different things *"tested in a browser"* means |
|
package/SKILL-CARD.md
CHANGED
|
@@ -12,7 +12,7 @@ harmless.
|
|
|
12
12
|
|---|---|
|
|
13
13
|
| **Purpose** | Runs a substantial task through ten gated delivery stages — intake grill, docs study, brainstorm, spec, plan, subagent build, tests, lint/deploy, post-deploy, docs+registers, acceptance — refusing to advance until each gate passes |
|
|
14
14
|
| **Owner** | ssheleg ([github.com/ssheleg/task-pipeline](https://github.com/ssheleg/task-pipeline)) |
|
|
15
|
-
| **Version** | 1.
|
|
15
|
+
| **Version** | 1.89.1 |
|
|
16
16
|
| **Surface** | Claude Code (filesystem skill + plugin) and the vercel `skills` CLI. **Not** uploaded to the Skills API; custom Skills do not sync across surfaces |
|
|
17
17
|
| **Dependencies** | None required. Optional: `context7` (MCP), `figma` (MCP), super-ux, agent-sync, graphify, obsidian-wiki, and **one of two browser channels** — `playwright` (CLI or MCP) or `chrome-devtools` (MCP); either satisfies the browser step and neither is required. Every stage's doctrine ships in-repo; the one conditional requirement is super-ux for the stage-3 UX track on a user-facing task |
|
|
18
18
|
| **Evaluation status** | Suite authored, 5 categories. One recorded run, **self-observed by the author**; **zero blind runs on zero of three models** — the split, and the numbers, live in [`evals/RESULTS.md`](evals/RESULTS.md) and are computed by `evals/run.py` |
|
package/package.json
CHANGED
|
@@ -1,13 +1,13 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "task-pipeline-skill",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.89.1",
|
|
4
4
|
"description": "Full-cycle delivery pipeline for coding agents: a mandatory built-in intake grill, then 10 gated stages (docs, brainstorm+decompose, spec, plan, build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine ships inside the skill — no companion plugin required. This package is the installer CLI.",
|
|
5
5
|
"bin": {
|
|
6
6
|
"task-pipeline": "bin/task-pipeline.js"
|
|
7
7
|
},
|
|
8
8
|
"scripts": {
|
|
9
|
-
"test": "python3 test/validate.py && python3 test/graph_test.py && python3 test/project_audit_test.py && python3 test/packet_schema_test.py && python3 test/browser_claims_test.py && python3 test/stage_checkpoint_test.py && npm run test:audit-regressions",
|
|
10
|
-
"test:all": "python3 test/validate.py && python3 test/graph_test.py && python3 test/project_audit_test.py && python3 test/packet_schema_test.py && python3 test/browser_claims_test.py && python3 test/stage_checkpoint_test.py && python3 test/negatives.py && npm run test:certify && npm run test:exposure && npm run test:probe && npm run test:anchors && npm run test:runner && npm run test:hooks && npm run test:artifacts && npm run test:docs && npm run test:audit-regressions",
|
|
9
|
+
"test": "python3 test/validate.py && python3 test/graph_test.py && python3 test/project_audit_test.py && python3 test/packet_schema_test.py && python3 test/browser_claims_test.py && python3 test/visual_gate_test.py && python3 test/stage_checkpoint_test.py && npm run test:audit-regressions",
|
|
10
|
+
"test:all": "python3 test/validate.py && python3 test/graph_test.py && python3 test/project_audit_test.py && python3 test/packet_schema_test.py && python3 test/browser_claims_test.py && python3 test/visual_gate_test.py && python3 test/stage_checkpoint_test.py && python3 test/negatives.py && npm run test:certify && npm run test:exposure && npm run test:probe && npm run test:anchors && npm run test:runner && npm run test:hooks && npm run test:artifacts && npm run test:docs && npm run test:audit-regressions",
|
|
11
11
|
"test:negatives": "python3 test/negatives.py",
|
|
12
12
|
"test:exposure": "python3 test/exposure_test.py",
|
|
13
13
|
"test:probe": "python3 test/probe.py --self-test",
|
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
"name": "task-pipeline",
|
|
4
4
|
"displayName": "Task Pipeline",
|
|
5
5
|
"description": "Runs a substantial task through a mandatory built-in intake grill, then 10 gated stages (docs, brainstorm+decompose, spec, plan, subagent build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine is built into the skill — no companion plugin required — with typed auto/judgment/manual gates, a frozen requirement spine that closes with evidence, a work board and a verification ledger that outlive a run, an exposure line naming what shipped unconfirmed, a progress rail computed from the project's own config, a loop guard whose review ceiling measures rather than stops, and stage-3 tracks for what a product does, how it sounds and how it looks. Two modes need no task: `checkup` (what is unverified) and `setup` (audit existing docs). Retro insights can publish upstream as issues, opt-in and redacted.",
|
|
6
|
-
"version": "1.
|
|
6
|
+
"version": "1.89.1",
|
|
7
7
|
"author": {
|
|
8
8
|
"name": "ssheleg",
|
|
9
9
|
"url": "https://x.com/sshlg93"
|
|
@@ -16,7 +16,9 @@ Your subject is **behaviour and what claims it.** Documentation, scenarios, the
|
|
|
16
16
|
changelog, ADRs, the runbook, the strings a user reads, and the other features that
|
|
17
17
|
share this path. You may open code to confirm a behaviour — but code is your
|
|
18
18
|
evidence, never your scope. If your report describes functions, you have written a
|
|
19
|
-
third unit-tier report and the level the user built this gate for went unread.
|
|
19
|
+
third unit-tier report and the level the user built this gate for went unread. The
|
|
20
|
+
rendered pixels are not yours either: on a flagship or product surface a fourth,
|
|
21
|
+
visual reading takes the contact sheet.
|
|
20
22
|
|
|
21
23
|
## Read the claims before you judge the change
|
|
22
24
|
|
|
@@ -0,0 +1,115 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: verifier-visual
|
|
3
|
+
description: Tier 4 of a task-pipeline certification, owed on a flagship or product surface. Reads the contact sheet, the director record, the project linter's output and the rubric, and reports whether what a user SEES carries the intent the record set — every state in the matrix, no gate item failing under a pass, no judge verdict an uncalibrated judge invented. Returns an eight-key tier report. Use as the fourth blind reading when a node whose surface_class is flagship or product claims to be finished. Not for code, call graphs or documentation — those are the unit, seam and product tiers.
|
|
4
|
+
model: inherit
|
|
5
|
+
tools: Read, Grep, Glob, Bash
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
# Visual tier — what a user sees, read against what was intended
|
|
9
|
+
|
|
10
|
+
You are the **fourth** reading, and the only one that opens a picture. Three agents are
|
|
11
|
+
reading the same node — the diff, what reaches it, what the product says about it — and
|
|
12
|
+
you will never see their reports, nor they yours. Blind is the point: an agent that has
|
|
13
|
+
just read *the implementation is sound* will look at a frame and see a sound
|
|
14
|
+
implementation.
|
|
15
|
+
|
|
16
|
+
Your subject is **the rendered surface against its stated intent.** Four inputs, and
|
|
17
|
+
nothing else is yours:
|
|
18
|
+
|
|
19
|
+
- the **contact sheet** — a filled `browser-claims/1` file whose look rows carry
|
|
20
|
+
`axes`, `capture`, `figma_frame`, `baseline`, `diff` and `rubric`, and the frames it
|
|
21
|
+
names;
|
|
22
|
+
- the **director record** — `docs/design/<surface>/director-record.md`: the brief and
|
|
23
|
+
its falsifier, the rubric written before any render, the signature moment, the
|
|
24
|
+
markers run;
|
|
25
|
+
- the **project linter's output** — `scripts/visual_gate.py lint`, or its NOT_RUN;
|
|
26
|
+
- the **rubric** — the items the record and the sheet name, each typed G (deterministic
|
|
27
|
+
gate), J (judge) or H (the person on the sheet).
|
|
28
|
+
|
|
29
|
+
If you find yourself reading a component's source to decide something, stop — that
|
|
30
|
+
finding belongs to the unit tier. If you are checking whether the README describes the
|
|
31
|
+
screen, that is the product tier's.
|
|
32
|
+
|
|
33
|
+
## What you actually do
|
|
34
|
+
|
|
35
|
+
1. **Run the deterministic checks first, and quote them.**
|
|
36
|
+
`python3 scripts/visual_gate.py sheet <sheet> --class <surface_class> --artifact-root
|
|
37
|
+
<frames>` and, where the record exists, `python3 scripts/visual_gate.py record <record>
|
|
38
|
+
--class <surface_class>`. Their output is your first `evidence` rows. A FAIL there is
|
|
39
|
+
a `breaks` finding, whatever the frames look like to you; a NOT_RUN is said as
|
|
40
|
+
NOT_RUN, never smoothed into a pass.
|
|
41
|
+
2. **Read the record before any frame.** The falsifier, the signature moment, the rubric.
|
|
42
|
+
The standard is what the record committed to before the render, not what the render
|
|
43
|
+
turned out to be.
|
|
44
|
+
3. **Walk the matrix, state by state.** For each `SCR` state: does a frame exist at the
|
|
45
|
+
mandatory pairs (dark × large text, RTL × narrow where the product has RTL), is its
|
|
46
|
+
capture record for this revision, does it diff clean against its Figma frame or
|
|
47
|
+
approved baseline? A hole or a stale frame is a finding about the sheet, not a
|
|
48
|
+
judgement about taste.
|
|
49
|
+
4. **Answer the rubric as a checklist — never as a score.** Each item is binary against
|
|
50
|
+
the record and the frame. Where you compare, compare **only against the approved
|
|
51
|
+
reference, in both orders**, and take **three samples**. A verdict that flips with the
|
|
52
|
+
order or between samples is **`uncertain`**: it goes in `not_examined` as
|
|
53
|
+
`uncertain — <item> — to the person on the sheet`, not into `confirms` and not into
|
|
54
|
+
`findings`.
|
|
55
|
+
5. **Mark every judge item NOT_ASSESSED until a labelled set calibrates it.** A J item
|
|
56
|
+
whose sheet row carries no `calibration` is not yours to pass or fail: list it in
|
|
57
|
+
`not_examined` as `NOT_ASSESSED — <item> — no labelled set`. A verdict from an
|
|
58
|
+
uncalibrated judge is a guess with a format.
|
|
59
|
+
6. **Never override a deterministic FAIL.** A G item or a linter S1 that failed is
|
|
60
|
+
`breaks`. You may add a finding beside it; you may not argue it away.
|
|
61
|
+
7. **Say what you read.** `scope` lists the sheet, the record, the frames you opened and
|
|
62
|
+
the linter output, each with its path. A pass on an empty `scope` is refused by
|
|
63
|
+
`graph.py certify` by name.
|
|
64
|
+
|
|
65
|
+
## `breaks` or `risk`, and the line is not taste
|
|
66
|
+
|
|
67
|
+
- **`breaks`** — the surface contradicts the record (the falsifier fails, the signature
|
|
68
|
+
moment is absent, a rubric item the record wrote first is violated on a frame), a G
|
|
69
|
+
item or an S1 failed, a required state has no frame, or a frame is stale or of the
|
|
70
|
+
wrong state. Every `breaks` carries a `check` — the command, or the judgement named as
|
|
71
|
+
one ("judgement — SCR-02/empty at 375 × dark × 200 %, against R-item and the record's
|
|
72
|
+
falsifier") — and its `fix` is a **triple: region → defect → fix**. A finding nobody
|
|
73
|
+
can locate on the frame is an impression.
|
|
74
|
+
- **`risk`** — found and survivable: a pairwise hole on a product surface, a frame whose
|
|
75
|
+
diff was NOT_RUN with a reason, an item the person should look at first. It ships,
|
|
76
|
+
named, as a blocker the run can continue around.
|
|
77
|
+
|
|
78
|
+
**A clean render is a valid result.** There is no quota of findings; inventing a defect
|
|
79
|
+
to look thorough starts a re-render round nobody needed, and the budget is one, two at
|
|
80
|
+
most.
|
|
81
|
+
|
|
82
|
+
## The report — all eight keys, and `[]` is an answer
|
|
83
|
+
|
|
84
|
+
```json
|
|
85
|
+
{
|
|
86
|
+
"node": "N-012",
|
|
87
|
+
"tier": "visual",
|
|
88
|
+
"verdict": "fail",
|
|
89
|
+
"scope": ["design/review/contact-sheet.json — 9 frames, SCR-01 and SCR-02",
|
|
90
|
+
"docs/design/landing/director-record.md — Brief, Rubric, Signature",
|
|
91
|
+
"design/review/lint.json — the project linter at 4dbb96c"],
|
|
92
|
+
"confirms": ["the falsifier holds on every 375-wide frame: the price is above the fold"],
|
|
93
|
+
"findings": [{ "what": "the empty state's helper text is grey on grey in dark at 200 % text",
|
|
94
|
+
"where": "contact-sheet.json BC-06 — SCR-02/empty, 375x667 · dark · 200% · ar",
|
|
95
|
+
"severity": "breaks",
|
|
96
|
+
"fix": "region: empty-state helper → defect: contrast 2.9:1 → fix: --sem-fg-muted to --sem-fg-default",
|
|
97
|
+
"check": "python3 scripts/visual_gate.py sheet design/review/contact-sheet.json --class product" }],
|
|
98
|
+
"evidence": ["visual_gate.py sheet … --class product → FAIL, 1 problem (BC-06 G item FAIL under PASS)",
|
|
99
|
+
"visual_gate.py record … --class product → PASS · validator NOT_RUN (sheleg-design without --check-record)"],
|
|
100
|
+
"not_examined": ["NOT_ASSESSED — R7 (J) — no labelled set calibrates the judge",
|
|
101
|
+
"uncertain — R12 on BC-08: the order-swapped comparison disagreed — to the person on the sheet"]
|
|
102
|
+
}
|
|
103
|
+
```
|
|
104
|
+
|
|
105
|
+
## Three ways this goes wrong
|
|
106
|
+
|
|
107
|
+
| Temptation | Why it is wrong |
|
|
108
|
+
|---|---|
|
|
109
|
+
| «It looks great» | That is a score with no checklist under it. Name the items, the frames and the reference, or name nothing |
|
|
110
|
+
| «The linter flagged it, but visually it's fine» | You sit below the deterministic floor. A G item or S1 that failed is `breaks`, and the fix is the code or the rule, not your opinion |
|
|
111
|
+
| «One more re-render will get it» | The budget is one, two at most. Past it, the item is `unresolved` on the sheet and goes to the person on their one pass |
|
|
112
|
+
|
|
113
|
+
Doctrine: `references/certification.md` → *The fourth reading*. The checks and the
|
|
114
|
+
contact sheet: `references/browser.md` → *The visual half*. The record's fields:
|
|
115
|
+
`references/spec.md` → *The COPY and VISUAL tracks*.
|
|
@@ -9,7 +9,8 @@ tools: Read, Grep, Glob, Bash
|
|
|
9
9
|
|
|
10
10
|
> **A node is normally closed by three readings, not by this one.**
|
|
11
11
|
> `verifier-unit`, `verifier-seam` and `verifier-product` each report at a
|
|
12
|
-
> different distance
|
|
12
|
+
> different distance — plus `verifier-visual` on a node whose `surface_class` is
|
|
13
|
+
> flagship or product — `graph.py certify` requires every owed tier to pass and assembles
|
|
13
14
|
> the verdict below from them — because a change can be correct where it was made
|
|
14
15
|
> and wrong one level out, and a single context cannot see both. Doctrine:
|
|
15
16
|
> `references/certification.md`. This agent remains for the case `certify` does not
|
|
@@ -135,7 +135,7 @@ Three things the grill does beyond clarifying the request, each in full in
|
|
|
135
135
|
[`references/grill.md`](references/grill.md):
|
|
136
136
|
- **Domain awareness** — it reads the project's `CONTEXT.md` / `docs/adr/` and holds the operator to them, writing resolved terms back as they land.
|
|
137
137
|
- **The autonomy sweep** — it pre-resolves what would otherwise stop stages 1→10 mid-flight. Autonomy is bought here or not at all; an unasked question is a scheduled interruption.
|
|
138
|
-
- **The design destination** with Figma on — *which* file, in which team, decided at stage 0. Left to drawing time it is answered by whoever holds the brush, and the answer is usually *create a new file*.
|
|
138
|
+
- **The design destination** with Figma on — *which* file per surface, in which team, decided at stage 0. Left to drawing time it is answered by whoever holds the brush, and the answer is usually *create a new file*.
|
|
139
139
|
|
|
140
140
|
## How to run
|
|
141
141
|
|
|
@@ -207,13 +207,13 @@ capable available — see `references/model-tiering.md`).
|
|
|
207
207
|
|
|
208
208
|
| # | Stage | Gate | Type |
|
|
209
209
|
|---|---|---|---|
|
|
210
|
-
| 0 | Intake grill — **mandatory** | source ledger written with its `Contradictions:` line; `docs/DOCMAP.md` answered and intent reconciled against as-built; the retro read in full; autonomy sweep covered; brief locked and confirmed | manual |
|
|
210
|
+
| 0 | Intake grill — **mandatory** | source ledger written with its `Contradictions:` line; `docs/DOCMAP.md` answered and intent reconciled against as-built; the retro read in full; autonomy sweep covered; UI: `surface_class` recorded; brief locked and confirmed | manual |
|
|
211
211
|
| 1 | Docs study | contracts grounded on fetched docs | auto |
|
|
212
212
|
| 2 | Brainstorm + decompose | design approved; UI verdict recorded; every REQ answered; **the queue is an artifact** — a work graph validates and its coverage names no unserved REQ; platform: module map approved | manual |
|
|
213
|
-
| 3 | Spec | committed + reviewed; UI: chain validated, linter green, scenarios and `SCR-` traced; COPY and VISUAL are a parallel layer after UX, and where both ran their convergence check is recorded | manual |
|
|
213
|
+
| 3 | Spec | committed + reviewed; UI: chain validated, linter green, scenarios and `SCR-` traced; COPY and VISUAL are a parallel layer after UX, and where both ran their convergence check is recorded; VISUAL is checked by its director record | manual |
|
|
214
214
|
| 4 | Plan | parallel-ready, DoD per task; **every edge names what it carries** — the fake-edge test run and its `Edges:` count computed; **`scripts/plan_audit.py` clean** — every live node's packet survives a cold reader, and no two unordered nodes edit one file | auto |
|
|
215
215
|
| 5 | Dev | tasks DONE, TDD green per task, branch integrated per the brief; a fanned-out group gets **one convergence check over all its diffs together** before the first worktree lands | auto |
|
|
216
|
-
| 6 | Tests | full suite green, new and changed code covered, every new check probed both ways and asserted on its exit code; **a web surface is checked in a browser, not in the diff** — where a browser channel is connected; absent, the weaker claim is recorded | auto |
|
|
216
|
+
| 6 | Tests | full suite green, new and changed code covered, every new check probed both ways and asserted on its exit code; **a web surface is checked in a browser, not in the diff** — where a browser channel is connected; absent, the weaker claim is recorded; a gated surface's contact sheet passes | auto |
|
|
217
217
|
| 7 | Lint + deploy | lint clean and suite green before deploy; deploy needs a go, or the brief's specific standing authorization | manual |
|
|
218
218
|
| 8 | Post-deploy | clean boot or an honest degradation report; **a deployed web target is opened, not curled** — a `200` proves the server answered and nothing else; where no browser channel is connected, the weaker claim is recorded | auto |
|
|
219
219
|
| 9 | Docs + wiki | every stale row of the stage-0 source ledger updated; the propagation matrix walked for every change type this run produced; the documentation gate green with its ratchets printed; docs, wiki and the code graph synced and checked against each other | auto |
|
|
@@ -196,9 +196,18 @@
|
|
|
196
196
|
},
|
|
197
197
|
"description": "What this node MUTATES — paths, register names, remote resource ids. `references/planning.md` states the rule the frontier needs: *distinct is not the same as independent, and the check is what they touch, never what they are called.* That rule lived in the markdown plan, and the graph replaced the plan as the thing deciding what runs next — so `next` could hand two agents two runnable nodes that write the same file, with nothing able to report it.\\n\\nOptional, and its absence is DISCLOSED rather than treated as «touches nothing»: `next` prints how many frontier nodes declared no targets, because a quiet run and a checked one must not look alike."
|
|
198
198
|
},
|
|
199
|
+
"surface_class": {
|
|
200
|
+
"enum": [
|
|
201
|
+
"flagship",
|
|
202
|
+
"product",
|
|
203
|
+
"internal",
|
|
204
|
+
"ad"
|
|
205
|
+
],
|
|
206
|
+
"description": "The class of the user-facing surface this node builds, copied from the brief's `surface_class` (stage 0). It selects the gate profile, and `graph.py certify` reads it: on `flagship` or `product` the fourth, `visual` tier report is REQUIRED, because none of unit, seam or product opens the contact sheet; on any other class a `visual` report is accepted when given. Absent on a node that builds no user-facing surface. `references/certification.md` → *The fourth reading*."
|
|
207
|
+
},
|
|
199
208
|
"certification": {
|
|
200
209
|
"type": "object",
|
|
201
|
-
"description": "What the
|
|
210
|
+
"description": "What the tiered certification recorded for this node. Written by `graph.py certify` on every round, pass or fail: a failing round that wrote nothing would erase the only evidence that a node is churning, which is the number the loop-guard ceiling reads.",
|
|
202
211
|
"additionalProperties": false,
|
|
203
212
|
"required": [
|
|
204
213
|
"round",
|
|
@@ -237,6 +246,12 @@
|
|
|
237
246
|
"pass",
|
|
238
247
|
"fail"
|
|
239
248
|
]
|
|
249
|
+
},
|
|
250
|
+
"visual": {
|
|
251
|
+
"enum": [
|
|
252
|
+
"pass",
|
|
253
|
+
"fail"
|
|
254
|
+
]
|
|
240
255
|
}
|
|
241
256
|
}
|
|
242
257
|
},
|
|
@@ -273,6 +288,12 @@
|
|
|
273
288
|
"pass",
|
|
274
289
|
"fail"
|
|
275
290
|
]
|
|
291
|
+
},
|
|
292
|
+
"visual": {
|
|
293
|
+
"enum": [
|
|
294
|
+
"pass",
|
|
295
|
+
"fail"
|
|
296
|
+
]
|
|
276
297
|
}
|
|
277
298
|
}
|
|
278
299
|
}
|
|
@@ -18,7 +18,7 @@
|
|
|
18
18
|
],
|
|
19
19
|
"gate": {
|
|
20
20
|
"type": "manual",
|
|
21
|
-
"check": "MANDATORY stage — never skipped (only sanctioned bypass: the entry-from-super-ux short-circuit). PHASE 1, before the first question: harvest the knowledge sources (references/knowledge-sources.md) — code, THE CODE GRAPH when one is built (references/knowledge-graph.md: graphify query/affected/god-nodes answer reach, which grep cannot; detect graphify-out/graph.json — recommended, never required), CLAUDE.md/AGENTS.md, CONTEXT.md + docs/adr, docs/ + docs/ux, past pipeline briefs and carry-over ledgers, THE RETRO'S STANDING INSTRUCTIONS (docs/evidence/retro.md — read IN FULL, not queried: they are capped at ten and they BIND this run; stamp each one the moment it fires, since that date is the only evidence behind stage 10's cold-retirement rule — references/retrospective.md), the knowledge wiki when installed (obsidian-wiki — recommended, never required; detect ~/.obsidian-wiki/config), and any other repo or hosted doc system the project names as its docs — queried by this task's own terms, with the SOURCE LEDGER written into the brief (a row per source consulted, or an explicit 'none found'; THE GRAPH'S ROW CARRIES ITS MEASURED LAG — commits and days behind HEAD, the signal it was measured with (built_at_commit exact / mtime approximate / unresolvable), and the marker '⚠ not trusted for reach until refreshed' — that exact string, so one marker is greppable across every ledger — on anything but current. A bare build date does NOT satisfy this: it is the graph's own reply about itself, true and self-reported and silent about whether the graph describes the tree this run is about to change — references/knowledge-graph.md -> Measure the lag, and references/gates.md -> False success for the class). PHASE 2, the grill, built into the skill (references/grill.md) — no companion to install. Per its contract: one question at a time, a recommended answer with each, explore the codebase/docs before asking, depth-first, contradictions reconciled; EVERY answer that touches a harvested source is validated against that source — the operator outranks any document, but only out loud, and the losing side is logged for the stage-9 doc update; domain awareness applied (terms challenged against CONTEXT.md, ADRs recorded for hard-to-reverse calls). The autonomy sweep is covered — every stage 1-10 has its blockers pre-resolved (docs sources, branch/tracker policy, test + lint commands, deploy target and authorization, log/health locations, docs+wiki targets, and for UI tasks the design surface: Figma on or text-only, is the Figma MCP connected, and if it is not — ship text-only or stop and connect it, since the UX chain degrades on its own and never blocks; plus, with Figma on, the DESIGN DESTINATION — which team/org by name and which file (the recorded one, a URL the operator gives, or creation in that named team with the creation explicitly authorized), written into the project's canonical record before the first frame, and never created while a recorded file resolves — an unreachable recorded file means stop and ask, never make a replacement) or is explicitly marked 'stop and ask here'. UI verdict recorded (arms super-ux); model decision recorded. All of it locked into a committed task brief the operator confirms before stage 1. The REQ table is written — one row per independently verifiable deliverable, each naming how it is verified — and frozen: adding later is free, removing or narrowing needs the operator's explicit agreement. The carry-over ledger is seeded. PHASE 1b, THE DOCUMENTATION INVENTORY (references/documentation.md): four questions answered into docs/DOCMAP.md before the interview — where settled things live (the DECISION HOME, and there is exactly one per project: an existing docs/adr/ IS the register and is recorded as such, never duplicated), what each fact's single home is, what a change of type X obliges (THE PROPAGATION MATRIX, non-empty, every row naming the check that enforces it or the word 'review' with a one-line reason), and what proves it (the gate command). A project with no answers gets them seeded — registers, matrix and scripts/check-docs.sh from the skill's templates — and the seeding is recorded as the register's first entry; the seeded gate must exit 0 on its own seeds, because a project that starts red teaches everyone on day one that the gate is noise. The regime is recorded. PHASE 1c, RECONCILE: git says how it should be, the run record says how it turned out — read both for the area about to be touched and resolve every divergence (the document is stale, the record is wrong, or they genuinely disagree and that is a decision), because starting on an unresolved divergence means building against a system that does not exist. The retro's in-force sections are read IN FULL and its archive is QUERIED by the task's nouns."
|
|
21
|
+
"check": "MANDATORY stage — never skipped (only sanctioned bypass: the entry-from-super-ux short-circuit). PHASE 1, before the first question: harvest the knowledge sources (references/knowledge-sources.md) — code, THE CODE GRAPH when one is built (references/knowledge-graph.md: graphify query/affected/god-nodes answer reach, which grep cannot; detect graphify-out/graph.json — recommended, never required), CLAUDE.md/AGENTS.md, CONTEXT.md + docs/adr, docs/ + docs/ux, past pipeline briefs and carry-over ledgers, THE RETRO'S STANDING INSTRUCTIONS (docs/evidence/retro.md — read IN FULL, not queried: they are capped at ten and they BIND this run; stamp each one the moment it fires, since that date is the only evidence behind stage 10's cold-retirement rule — references/retrospective.md), the knowledge wiki when installed (obsidian-wiki — recommended, never required; detect ~/.obsidian-wiki/config), and any other repo or hosted doc system the project names as its docs — queried by this task's own terms, with the SOURCE LEDGER written into the brief (a row per source consulted, or an explicit 'none found'; THE GRAPH'S ROW CARRIES ITS MEASURED LAG — commits and days behind HEAD, the signal it was measured with (built_at_commit exact / mtime approximate / unresolvable), and the marker '⚠ not trusted for reach until refreshed' — that exact string, so one marker is greppable across every ledger — on anything but current. A bare build date does NOT satisfy this: it is the graph's own reply about itself, true and self-reported and silent about whether the graph describes the tree this run is about to change — references/knowledge-graph.md -> Measure the lag, and references/gates.md -> False success for the class). PHASE 2, the grill, built into the skill (references/grill.md) — no companion to install. Per its contract: one question at a time, a recommended answer with each, explore the codebase/docs before asking, depth-first, contradictions reconciled; EVERY answer that touches a harvested source is validated against that source — the operator outranks any document, but only out loud, and the losing side is logged for the stage-9 doc update; domain awareness applied (terms challenged against CONTEXT.md, ADRs recorded for hard-to-reverse calls). The autonomy sweep is covered — every stage 1-10 has its blockers pre-resolved (docs sources, branch/tracker policy, test + lint commands, deploy target and authorization, log/health locations, docs+wiki targets, and for UI tasks the design surface: Figma on or text-only, is the Figma MCP connected, and if it is not — ship text-only or stop and connect it, since the UX chain degrades on its own and never blocks; plus, with Figma on, the DESIGN DESTINATION — which team/org by name and which file per surface, App / Web / ASO (the recorded one, a URL the operator gives, or creation in that named team with the creation explicitly authorized), written into the project's canonical record before the first frame, and never created while a recorded file resolves — an unreachable recorded file means stop and ask, never make a replacement) or is explicitly marked 'stop and ask here'. UI verdict recorded (arms super-ux), and for a user-facing task its SURFACE CLASS — surface_class: flagship | product | internal | ad — which selects the visual gate profile for stages 3, 6 and 10 (references/stages.md -> The surface class); model decision recorded. All of it locked into a committed task brief the operator confirms before stage 1. The REQ table is written — one row per independently verifiable deliverable, each naming how it is verified — and frozen: adding later is free, removing or narrowing needs the operator's explicit agreement. The carry-over ledger is seeded. PHASE 1b, THE DOCUMENTATION INVENTORY (references/documentation.md): four questions answered into docs/DOCMAP.md before the interview — where settled things live (the DECISION HOME, and there is exactly one per project: an existing docs/adr/ IS the register and is recorded as such, never duplicated), what each fact's single home is, what a change of type X obliges (THE PROPAGATION MATRIX, non-empty, every row naming the check that enforces it or the word 'review' with a one-line reason), and what proves it (the gate command). A project with no answers gets them seeded — registers, matrix and scripts/check-docs.sh from the skill's templates — and the seeding is recorded as the register's first entry; the seeded gate must exit 0 on its own seeds, because a project that starts red teaches everyone on day one that the gate is noise. The regime is recorded. PHASE 1c, RECONCILE: git says how it should be, the run record says how it turned out — read both for the area about to be touched and resolve every divergence (the document is stale, the record is wrong, or they genuinely disagree and that is a decision), because starting on an unresolved divergence means building against a system that does not exist. The retro's in-force sections are read IN FULL and its archive is QUERIED by the task's nouns."
|
|
22
22
|
}
|
|
23
23
|
},
|
|
24
24
|
{
|
|
@@ -65,7 +65,7 @@
|
|
|
65
65
|
],
|
|
66
66
|
"gate": {
|
|
67
67
|
"type": "manual",
|
|
68
|
-
"check": "UX track ran FIRST for user-facing tasks (/ux -> ux-foundation CJM -> ux-flows screens -> ux-scenarios -> /ux-lint green); spec committed and user-reviewed; every user-facing requirement traces to a scenario ID. Every spec section carries covers: REQ-... and every REQ appears in at least one section. With Figma on: the destination the brief named was used — the canonical record (docs/ux/foundation.md -> Design tooling)
|
|
68
|
+
"check": "UX track ran FIRST for user-facing tasks (/ux -> ux-foundation CJM -> ux-flows screens -> ux-scenarios -> /ux-lint green); spec committed and user-reviewed; every user-facing requirement traces to a scenario ID. Every spec section carries covers: REQ-... and every REQ appears in at least one section. With Figma on: the destination the brief named was used — the canonical record (docs/ux/foundation.md -> Design tooling) names one file per surface (App, Web, ASO), no file was created while a recorded one resolved, and every screens.md frame link's :fileKey is one of the recorded set (a string match, not a judgement — scripts/visual_gate.py filekeys; a key outside the set means the run drew in a file nobody recorded and nobody will open). COPY and VISUAL are a parallel layer after UX: every user-facing string went through the COPY track or the refusal is recorded, the visual layer went through the VISUAL track or the refusal is recorded — and on a flagship, product or ad surface the VISUAL track is checked by its TRACE, not by the fact it ran: the director record at docs/design/<surface>/director-record.md carries the fields its surface_class owes ('scripts/visual_gate.py record <file> --class <surface_class>' exits 0; sheleg-design's own --check-record validator runs where installed and reads NOT_RUN, never PASS, where it is not; 'Mode: declined' with its reason is the refusal and passes), and where both tracks ran their convergence check is recorded - findings with the ruling, or 'Tracks converge: clean'."
|
|
69
69
|
}
|
|
70
70
|
},
|
|
71
71
|
{
|
|
@@ -106,7 +106,7 @@
|
|
|
106
106
|
],
|
|
107
107
|
"gate": {
|
|
108
108
|
"type": "auto",
|
|
109
|
-
"check": "full suite green (not just new tests); new/changed code covered including failure paths; no skip/xfail smuggling a red suite past the gate; tests assert real behavior, not mock behavior",
|
|
109
|
+
"check": "full suite green (not just new tests); new/changed code covered including failure paths; no skip/xfail smuggling a red suite past the gate; tests assert real behavior, not mock behavior. Where the stage-3 VISUAL track ran on a flagship, product or ad surface, THE VISUAL HALF of the look passes too (references/browser.md -> The visual half): the contact sheet — frames over the SCR states x viewport x theme x text x locale, pairwise plus the mandatory pairs, each with its capture record, diffed against its Figma frame or approved baseline, the project linter run, judge items NOT_ASSESSED until calibrated — and 'scripts/visual_gate.py sheet <file> --class <surface_class>' exits 0 (NOT_RUN, exit 3, is not green); re-render budget one, two at most, then unresolved to the person. On internal it is recommended and the linter is the floor",
|
|
110
110
|
"command": "npm test"
|
|
111
111
|
}
|
|
112
112
|
},
|
|
@@ -167,7 +167,7 @@
|
|
|
167
167
|
],
|
|
168
168
|
"gate": {
|
|
169
169
|
"type": "manual",
|
|
170
|
-
"check": "Close the circle. FIRST the LADDER WALK (references/audit.md), because the REQ table can only find what was named and lost — a comparison needs two sides and an absence has one: walk each REQ bottom-up through its rungs (decision -> spec section -> contract AND its failure behavior -> plan task -> change -> executed test -> surface/docs), check the seam at each step, order findings BY SEAM not by file, and turn every absence into a new REQ row with its check BEFORE the table is written; findings belonging to a lower layer go back to that layer (spec -> stage 3, plan -> stage 4); record the pass's two counts (new findings vs findings caused by this run's own fixes) so the next pass can tell whether the axis is exhausted. THEN the coverage table: every REQ has a status (verified / partial / deferred / dropped) — none unknown; every verified carries evidence (a passing test name, file:line, a command and its output, or a scenario ID) — 'done' without evidence is downgraded to partial, not upgraded, and a green from a check nobody has watched fail against a planted defect is not evidence at all; every partial names what is missing and where it is tracked; every deferred/dropped has the operator's agreement and, for deferred, a tracker entry; no carry-over row is left unresolved and the ledger's counts are printed beside this verdict, so 'green' never reads as 'verified'; EVERY REPOSITORY IS CLOSED, THE PARENT INCLUDED — a submodule is finished only when its parent points at it, so 'git submodule status' shows no line starting with '+' and every repo is clean and pushed ('git -C <repo> status --porcelain' and 'git -C <repo> log @{u}..HEAD' both empty), because a parent records a submodule as a pointer to one commit and moving the submodule does not move the pointer: neither repo looks wrong alone and the disagreement survives every check that runs inside one; and the operator answers the closing question — here is what you asked for, here is what shipped, here is what is deferred, what is missing? — and signs off. LAST ACT, THE RETROSPECTIVE (references/retrospective.md, written to docs/evidence/retro.md — one file per project, not per run, because every gate in this flow is good at THIS run and blind across runs): STAMP THE RUN FIRST (date, topic, verdict, counts) — the order is load-bearing, not style: the cold-retirement trigger reads the stamp this stage writes, so a prune placed ahead of the stamp can never run on real data; THEN PRUNE — every standing instruction checked against its retirement triggers (it became a check; every path/command/stage it names is gone; it has not fired in the last five run stamps, or in the last sixty days), the list held to its hard cap of ten (at eleven the oldest never-fired row goes — 'they all matter' is the state in which the list stopped being read), and EVERY DELETION LOGGED as one line, never silent; THEN, only if the run diverged, write the entry — symptom with evidence, the stage it surfaced at, the stage that OWNED it, the root cause ('the agent was careless' is not one), the fix by grade (mechanical check > standing instruction with its retire-when written at birth > a note that expires in two runs), and the check that catches it the first time from now on. A retro left empty after a messy run is the failure this file exists to stop, and the retro counts are printed beside this gate's verdict like the carry-over ledger's, so a list that quietly grew back is visible where it happened. EVERY LESSON CARRIES ITS COMMIT: each standing instruction has the SHA that introduced it and the SHA of the run in which it last fired, each log entry and each retirement carries one, the run stamp carries the run's own — a file:line rots at the next edit while 'git show <sha>' reconstructs the whole incident two months later — and every SHA must resolve, which the documentation gate checks with 'git rev-parse --verify'. ROTATION: entries older than the last five run stamps MOVE into docs/evidence/retro/YYYY-QN.md, which is append-only and QUERIED rather than read, so the in-force file stays short enough to be read in full and pruning costs no knowledge. AND THE GATE ITSELF IS PROVEN: every check this close-out leans on — the documentation gate included — has been seen failing once against a planted defect, with the probe recorded, and its ratchet counts are printed beside this verdict (references/gates.md). THE HAND-BACK IS WRITTEN — the request quoted as GIVEN, progress against it, what was solved with evidence, what surfaced unasked, every waiting decision ASKED here with options, and the ambiguity count computed from the four registers; zero prints as zero."
|
|
170
|
+
"check": "Close the circle. FIRST the LADDER WALK (references/audit.md), because the REQ table can only find what was named and lost — a comparison needs two sides and an absence has one: walk each REQ bottom-up through its rungs (decision -> spec section -> contract AND its failure behavior -> plan task -> change -> executed test -> surface/docs), check the seam at each step — and where the VISUAL track ran, the row visual intent (the director record) <-> final render (the APPROVED contact sheet: 'visual_gate.py sheet ... --require-approval' exits 0 on flagship, product or ad), with the review rounds it took copied from the run ledger's review: lines into the acceptance file — order findings BY SEAM not by file, and turn every absence into a new REQ row with its check BEFORE the table is written; findings belonging to a lower layer go back to that layer (spec -> stage 3, plan -> stage 4); record the pass's two counts (new findings vs findings caused by this run's own fixes) so the next pass can tell whether the axis is exhausted. THEN the coverage table: every REQ has a status (verified / partial / deferred / dropped) — none unknown; every verified carries evidence (a passing test name, file:line, a command and its output, or a scenario ID) — 'done' without evidence is downgraded to partial, not upgraded, and a green from a check nobody has watched fail against a planted defect is not evidence at all; every partial names what is missing and where it is tracked; every deferred/dropped has the operator's agreement and, for deferred, a tracker entry; no carry-over row is left unresolved and the ledger's counts are printed beside this verdict, so 'green' never reads as 'verified'; EVERY REPOSITORY IS CLOSED, THE PARENT INCLUDED — a submodule is finished only when its parent points at it, so 'git submodule status' shows no line starting with '+' and every repo is clean and pushed ('git -C <repo> status --porcelain' and 'git -C <repo> log @{u}..HEAD' both empty), because a parent records a submodule as a pointer to one commit and moving the submodule does not move the pointer: neither repo looks wrong alone and the disagreement survives every check that runs inside one; and the operator answers the closing question — here is what you asked for, here is what shipped, here is what is deferred, what is missing? — and signs off. LAST ACT, THE RETROSPECTIVE (references/retrospective.md, written to docs/evidence/retro.md — one file per project, not per run, because every gate in this flow is good at THIS run and blind across runs): STAMP THE RUN FIRST (date, topic, verdict, counts) — the order is load-bearing, not style: the cold-retirement trigger reads the stamp this stage writes, so a prune placed ahead of the stamp can never run on real data; THEN PRUNE — every standing instruction checked against its retirement triggers (it became a check; every path/command/stage it names is gone; it has not fired in the last five run stamps, or in the last sixty days), the list held to its hard cap of ten (at eleven the oldest never-fired row goes — 'they all matter' is the state in which the list stopped being read), and EVERY DELETION LOGGED as one line, never silent; THEN, only if the run diverged, write the entry — symptom with evidence, the stage it surfaced at, the stage that OWNED it, the root cause ('the agent was careless' is not one), the fix by grade (mechanical check > standing instruction with its retire-when written at birth > a note that expires in two runs), and the check that catches it the first time from now on. A retro left empty after a messy run is the failure this file exists to stop, and the retro counts are printed beside this gate's verdict like the carry-over ledger's, so a list that quietly grew back is visible where it happened. EVERY LESSON CARRIES ITS COMMIT: each standing instruction has the SHA that introduced it and the SHA of the run in which it last fired, each log entry and each retirement carries one, the run stamp carries the run's own — a file:line rots at the next edit while 'git show <sha>' reconstructs the whole incident two months later — and every SHA must resolve, which the documentation gate checks with 'git rev-parse --verify'. ROTATION: entries older than the last five run stamps MOVE into docs/evidence/retro/YYYY-QN.md, which is append-only and QUERIED rather than read, so the in-force file stays short enough to be read in full and pruning costs no knowledge. AND THE GATE ITSELF IS PROVEN: every check this close-out leans on — the documentation gate included — has been seen failing once against a planted defect, with the probe recorded, and its ratchet counts are printed beside this verdict (references/gates.md). THE HAND-BACK IS WRITTEN — the request quoted as GIVEN, progress against it, what was solved with evidence, what surfaced unasked, every waiting decision ASKED here with options, and the ambiguity count computed from the four registers; zero prints as zero."
|
|
171
171
|
}
|
|
172
172
|
}
|
|
173
173
|
],
|
|
@@ -48,6 +48,16 @@ surface and docs), checking the seam at each step. It is one pass, scoped to thi
|
|
|
48
48
|
run's deliverables, and it is the only part of the pipeline that can find a gap
|
|
49
49
|
that was never a row.
|
|
50
50
|
|
|
51
|
+
**Where the stage-3 VISUAL track ran, the walk carries a row the REQ table cannot:
|
|
52
|
+
visual intent (the director record) ↔ final render (the approved contact sheet)** —
|
|
53
|
+
`audit.md`'s `V→R` seam. Read the record's falsifier, signature moment and rubric
|
|
54
|
+
against the sheet the person approved, frame by frame where they disagree. On a
|
|
55
|
+
`flagship`, `product` or `ad` surface the sheet must be approved — `python3
|
|
56
|
+
scripts/visual_gate.py sheet <file> --class <surface_class> --require-approval` exits
|
|
57
|
+
0 — and the rounds it took (the last `review:` line per surface in the run ledger) are
|
|
58
|
+
written into the acceptance file, because the ledger does not outlive the run and the
|
|
59
|
+
count of human passes is the number this whole layer exists to bring down.
|
|
60
|
+
|
|
51
61
|
- **An absence found here becomes a new REQ row with its check**, then the table is
|
|
52
62
|
written. The list is frozen against *narrowing*, never against additions
|
|
53
63
|
([`grill.md`](grill.md) → *The REQ spine*). Writing the table first and appending
|
|
@@ -69,6 +79,8 @@ Read all of them before writing anything:
|
|
|
69
79
|
- git log for the run's branch; the test suite's final output
|
|
70
80
|
- stage 8's post-deploy notes; stage 9's doc/wiki changes
|
|
71
81
|
- for UI tasks: `docs/ux/scenarios.md` statuses and the `/ux-lint` result
|
|
82
|
+
- where the VISUAL track ran: the director record, the approved contact sheet, and the
|
|
83
|
+
run ledger's `review:` lines
|
|
72
84
|
|
|
73
85
|
## Every declared stage is accounted for
|
|
74
86
|
|
|
@@ -110,7 +110,8 @@ absence findable.
|
|
|
110
110
|
| **L5** | Change | the commits — the thing actually in the tree |
|
|
111
111
|
| **L6** | Test | an **executed** assertion, by name — never "the tests pass" |
|
|
112
112
|
| **L7** | Surface | what a user reaches: scenario, screen state, CLI output, runbook |
|
|
113
|
-
| **F** | Frame — *conditional* | UI work with Figma on: one frame per `SCR-NN/<Screen>/<state>`, **in
|
|
113
|
+
| **F** | Frame — *conditional* | UI work with Figma on: one frame per `SCR-NN/<Screen>/<state>`, **in one of the files the project recorded — one per surface (App, Web, ASO)**. Not a step in the sequence — a **second, parallel statement of the same surface**, made in pictures |
|
|
114
|
+
| **V** | Visual intent — *conditional* | UI work where the stage-3 VISUAL track ran: the director record (`docs/design/<surface>/director-record.md`) — the falsifier, the signature moment, the rubric written before any render. Its counterpart **R** is the final render as the person approved it: the contact sheet (`browser.md` → *The visual half*) |
|
|
114
115
|
|
|
115
116
|
**Audit the seams, not the artifacts.** Each rung is internally consistent most of
|
|
116
117
|
the time — that is exactly what the horizontal pass is good at, and it has already
|
|
@@ -128,7 +129,8 @@ done it. What survives lives between rungs:
|
|
|
128
129
|
| L7→L0 | does the shipped surface satisfy the requirement's **statement**? | it does what the task said and not what the requirement meant |
|
|
129
130
|
| L2→F | *(UI)* does the frame render what the spec **says**? | a frame that promises a capability, limit or number the product does not have |
|
|
130
131
|
| F→L7 | *(UI)* did what shipped match the frame, or did the frame become fiction? | the frame is still the design of record and no longer describes anything that exists |
|
|
131
|
-
| →F | *(UI)* is every frame **in
|
|
132
|
+
| →F | *(UI)* is every frame **in a recorded file**? | a design file outside the recorded set — one per surface — that nobody opens, holding real work; the check is a `:fileKey` set membership (`visual_gate.py filekeys`), so it is a gate, not an opinion |
|
|
133
|
+
| V→R | *(UI, VISUAL track ran)* does the **approved contact sheet** carry the intent the **director record** set? | a render that passed every frame check and lost the falsifier, the signature moment or a rubric item the record wrote first; or a sheet nobody approved, standing in for the person's pass. The rounds it took (`review:` lines) go into the acceptance file |
|
|
132
134
|
|
|
133
135
|
The L7→L0 seam is stage 10's question, expressed as a seam. When it fails, the run
|
|
134
136
|
did every instruction correctly and delivered the wrong thing.
|
|
@@ -161,12 +163,16 @@ a deploy, or docs in another repository. **Creating** one is stronger still: it
|
|
|
161
163
|
needs a named team and an explicit authorization recorded at intake
|
|
162
164
|
([`grill.md`](grill.md) → *The design destination*).
|
|
163
165
|
|
|
164
|
-
**And check the
|
|
165
|
-
`figma.com/design/:fileKey/…`, so
|
|
166
|
-
the canonical record (`docs/ux/foundation.md` → *Design tooling
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
166
|
+
**And check the files, not just the frames.** Every deep link is
|
|
167
|
+
`figma.com/design/:fileKey/…`, so checking each `screens.md` link's key against
|
|
168
|
+
the set the canonical record names (`docs/ux/foundation.md` → *Design tooling*: one
|
|
169
|
+
file per surface — App, Web, ASO) is a string match —
|
|
170
|
+
`python3 scripts/visual_gate.py filekeys --record docs/ux/foundation.md --screens
|
|
171
|
+
docs/ux/screens.md`. A key outside the set is a **file nobody recorded, with real work
|
|
172
|
+
in it** — the failure that starts with one agent unable to open the recorded file and
|
|
173
|
+
quietly making a new one. One file per surface is the shape, not one file for every
|
|
174
|
+
frame: a store asset and an app screen have different owners, sizes and reviewers, and a
|
|
175
|
+
single file holding both is the file nobody can hand to either. Nothing else in the chain notices: the new file is internally consistent,
|
|
170
176
|
its frames are named correctly, and the linter is green.
|
|
171
177
|
|
|
172
178
|
## How one audit pass runs
|