@hecer/yoke 1.11.0 → 1.13.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +13 -13
- package/.codex-plugin/plugin.json +7 -7
- package/CHANGELOG.md +435 -398
- package/README.md +943 -915
- package/TODOS.md +5 -5
- package/agents/docs.toml +6 -6
- package/agents/implementer.toml +6 -6
- package/agents/reviewer.toml +6 -6
- package/agents/security.toml +6 -6
- package/bench/README.md +86 -86
- package/bench/RESULTS.md +35 -35
- package/bench/output-compaction.mjs +65 -65
- package/bench/result-schema.mjs +12 -12
- package/bench/results/claude-2026-07-27T18-03-26.json +50 -50
- package/bench/results/codex-unavailable-1785175418318.json +15 -15
- package/bench/results/gemini-2026-07-27T18-03-44.json +46 -46
- package/bench/run-matrix.mjs +26 -26
- package/bench/run.mjs +106 -106
- package/canon/AGENTS.md +30 -30
- package/canon/context/DECISIONS.md +4 -4
- package/canon/context/GLOSSARY.md +11 -11
- package/canon/context/KNOWLEDGE.md +4 -4
- package/canon/context/PROJECT.md +15 -15
- package/canon/loop/loop-spec.md +65 -65
- package/canon/loop/prd.schema.md +41 -41
- package/canon/manifest.yaml +59 -59
- package/canon/policy/gates.md +7 -7
- package/canon/policy/roles.md +9 -9
- package/canon/skills/ATTRIBUTION.md +99 -99
- package/canon/skills/authoring-prd/SKILL.md +57 -57
- package/canon/skills/brainstorming/SKILL.md +164 -164
- package/canon/skills/codebase-design/DEEPENING.md +15 -15
- package/canon/skills/codebase-design/DESIGN-IT-TWICE.md +12 -12
- package/canon/skills/codebase-design/SKILL.md +39 -39
- package/canon/skills/dispatching-parallel-agents/SKILL.md +182 -182
- package/canon/skills/document-release/SKILL.md +302 -302
- package/canon/skills/domain-modeling/ADR-FORMAT.md +19 -19
- package/canon/skills/domain-modeling/CONTEXT-FORMAT.md +39 -39
- package/canon/skills/domain-modeling/SKILL.md +35 -35
- package/canon/skills/executing-plans/SKILL.md +70 -70
- package/canon/skills/finishing-a-development-branch/SKILL.md +200 -200
- package/canon/skills/health/SKILL.md +177 -177
- package/canon/skills/maintaining-context/SKILL.md +34 -34
- package/canon/skills/minimal-code/SKILL.md +21 -21
- package/canon/skills/no-ai-slop/SKILL.md +103 -103
- package/canon/skills/no-ai-slop/eval.md +43 -43
- package/canon/skills/plan-ceo-review/SKILL.md +541 -541
- package/canon/skills/plan-eng-review/SKILL.md +362 -362
- package/canon/skills/receiving-code-review/SKILL.md +213 -213
- package/canon/skills/requesting-code-review/SKILL.md +105 -105
- package/canon/skills/resolving-merge-conflicts/SKILL.md +18 -18
- package/canon/skills/retro/SKILL.md +397 -397
- package/canon/skills/review/SKILL.md +246 -246
- package/canon/skills/ship/SKILL.md +691 -691
- package/canon/skills/subagent-driven-development/SKILL.md +277 -277
- package/canon/skills/systematic-debugging/SKILL.md +296 -296
- package/canon/skills/tdd/SKILL.md +371 -371
- package/canon/skills/unslop-ui/SKILL.md +34 -34
- package/canon/skills/using-git-worktrees/SKILL.md +218 -218
- package/canon/skills/verification-before-completion/SKILL.md +139 -139
- package/canon/skills/visual-verification/SKILL.md +54 -54
- package/canon/skills/workflow/SKILL.md +22 -22
- package/canon/skills/writing-for-agents/SKILL-MECHANICS.md +27 -27
- package/canon/skills/writing-for-agents/SKILL.md +42 -42
- package/canon/skills/writing-plans/SKILL.md +152 -152
- package/canon/skills/writing-skills/SKILL.md +655 -655
- package/canon/skills/yoke-retrofit/SKILL.md +26 -26
- package/canon/skills/yoke-workflow/SKILL.md +20 -20
- package/canon/tools/codex-rtk-hook.mjs +35 -35
- package/canon/tools/gemini-rtk-hook.mjs +25 -25
- package/canon/tools/graphify.md +3 -3
- package/canon/tools/playwright-mcp.md +3 -3
- package/canon/tools/qwen-rtk-hook.mjs +25 -0
- package/canon/tools/rtk.md +7 -7
- package/canon/tools/serena.md +6 -6
- package/dist/agents/catalog.js +7 -0
- package/dist/agents/contracts.js +3 -1
- package/dist/agents/host.js +5 -1
- package/dist/agents/process-streams.js +62 -0
- package/dist/agents/process.js +43 -3
- package/dist/agents/providers.js +61 -6
- package/dist/agents/telemetry.js +133 -37
- package/dist/canon/manifest.js +2 -1
- package/dist/change/inbox.js +1 -1
- package/dist/cli.js +30 -24
- package/dist/dashboard/page.js +122 -122
- package/dist/dashboard/panels.js +91 -91
- package/dist/goals/command.js +3 -2
- package/dist/loop/claims.js +2 -1
- package/dist/loop/decision.js +3 -2
- package/dist/loop/parallel-command.js +4 -2
- package/dist/loop/prd.js +2 -1
- package/dist/loop/reporter.js +1 -0
- package/dist/loop/run-command.js +31 -10
- package/dist/prd/command.js +19 -19
- package/dist/quality/candidate-comparison.js +6 -1
- package/dist/quality/command.js +17 -2
- package/dist/quality/types.js +6 -1
- package/dist/retrofit/apply.js +95 -2
- package/dist/retrofit/config.js +9 -1
- package/dist/retrofit/detect.js +8 -0
- package/dist/retrofit/plan.js +6 -0
- package/dist/retrofit/planners/claude.js +14 -14
- package/dist/retrofit/planners/kilo.js +44 -0
- package/dist/retrofit/planners/opencode.js +44 -0
- package/dist/retrofit/planners/pi.js +24 -0
- package/dist/retrofit/planners/qwen.js +3 -3
- package/dist/retrofit/preserve.js +2 -2
- package/dist/retrofit/qwen-settings.js +17 -0
- package/dist/retrofit/skill-actions.js +4 -1
- package/dist/retrofit/tools.js +8 -0
- package/dist/review/command.js +3 -2
- package/dist/review/verdict.js +1 -1
- package/dist/routing/capability.js +2 -2
- package/dist/routing/planning.js +2 -0
- package/dist/routing/registry.js +3 -1
- package/dist/routing/router.js +7 -3
- package/dist/setup/command.js +35 -11
- package/dist/setup/model-presets.js +48 -0
- package/docs/CAPABILITY-ROUTING.md +51 -51
- package/docs/DASHBOARD-EVOLUTION.md +33 -33
- package/docs/HARNESSES.md +81 -0
- package/docs/MIGRATING-TO-1.0.md +33 -33
- package/docs/MIGRATING-TO-1.1.md +27 -27
- package/docs/MIGRATING-TO-1.4.md +70 -70
- package/docs/PRODUCT-DIRECTION-2026-09-05.md +210 -210
- package/docs/PUBLISHING.md +114 -114
- package/docs/QWEN-MODEL-SUPPORT.md +142 -0
- package/docs/VERIFIED-PROJECTS-VALIDATION.md +29 -29
- package/docs/VERIFIED-PROJECTS.md +167 -167
- package/docs/superpowers/plans/2026-06-28-baustein-e-context-layer.md +981 -981
- package/docs/superpowers/plans/2026-06-29-baustein-f-routing.md +258 -258
- package/docs/superpowers/plans/2026-06-29-baustein-g-loop-observability.md +1006 -1006
- package/docs/superpowers/plans/2026-06-29-baustein-h-loop-robustness.md +374 -374
- package/docs/superpowers/plans/2026-06-30-baustein-i-visual-design-verification.md +450 -450
- package/docs/superpowers/plans/2026-07-02-baustein-k-zero-to-100-bootstrap.md +1024 -1024
- package/docs/superpowers/plans/2026-07-02-baustein-m-flow-smoke-proofs.md +574 -574
- package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -537
- package/docs/superpowers/plans/2026-08-16-artifact-backed-output-compaction.md +329 -329
- package/docs/superpowers/plans/2026-09-05-verified-projects.md +83 -83
- package/docs/superpowers/specs/2026-06-28-baustein-e-context-layer-design.md +146 -146
- package/docs/superpowers/specs/2026-06-29-baustein-f-routing-design.md +106 -106
- package/docs/superpowers/specs/2026-06-29-baustein-g-loop-observability-design.md +186 -186
- package/docs/superpowers/specs/2026-06-29-baustein-h-loop-robustness-design.md +113 -113
- package/docs/superpowers/specs/2026-06-30-baustein-i-visual-design-verification-design.md +98 -98
- package/docs/superpowers/specs/2026-07-02-baustein-k-zero-to-100-bootstrap-design.md +200 -200
- package/docs/superpowers/specs/2026-07-02-baustein-m-flow-smoke-proofs-design.md +155 -155
- package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -422
- package/docs/superpowers/specs/2026-08-16-artifact-backed-output-compaction-design.md +166 -166
- package/gemini-extension.json +6 -6
- package/hooks/hooks.json +19 -19
- package/package.json +91 -87
- package/dist/dashboard/discovery.js +0 -73
- package/docs/community-outreach-2026-08-20.md +0 -85
- package/docs/launch-copy-2026-08-21.md +0 -193
|
@@ -1,113 +1,113 @@
|
|
|
1
|
-
# Baustein H — Loop Robustness (verify-as-truth + flaky-tolerant verify)
|
|
2
|
-
|
|
3
|
-
**Status:** Design approved 2026-06-29
|
|
4
|
-
**Component:** Yoke (🐂)
|
|
5
|
-
**Relates to:** [[harness-build-progress]], [[harness-loop-technique]]
|
|
6
|
-
|
|
7
|
-
## Problem & Goal
|
|
8
|
-
|
|
9
|
-
Evidence from a real autonomous run (NewMarket, 52/57 stories): the loop blocked and needed
|
|
10
|
-
human intervention on a handful of stories. Two of those block causes are **harness defects**,
|
|
11
|
-
not project defects, and both are cheap to fix:
|
|
12
|
-
|
|
13
|
-
1. **Runner exit-code is treated as ground truth (SE5).** The story was implemented and the full
|
|
14
|
-
suite was green (769/769), but the `claude` `.cmd` wrapper on Windows exited **127**. The loop
|
|
15
|
-
checks `result.success` *before* running verify, so it **blocked a successful story** on a
|
|
16
|
-
spurious exit code and didn't auto-commit. This contradicts the loop's own stated philosophy
|
|
17
|
-
(Baustein C2): *"independent verification is the source of truth, not the agent's exit code."*
|
|
18
|
-
|
|
19
|
-
2. **No flaky-test tolerance (T4).** Under autonomous load, a heavy background task caused an
|
|
20
|
-
async test to time out — a **false** red, not a real defect. The single-shot verify gate
|
|
21
|
-
blocked the story. The user had to manually add `retry: 2` to the project's `vitest.config`.
|
|
22
|
-
That resilience belongs in the harness.
|
|
23
|
-
|
|
24
|
-
**Goal:** make the loop trust **verify**, not the agent's exit code, and tolerate transient
|
|
25
|
-
flakes — closing two real block causes without weakening the gate.
|
|
26
|
-
|
|
27
|
-
**Out of scope (deferred):** cross-story regression repair (S5/S6 — a later story breaking an
|
|
28
|
-
earlier story's tests). That needs its own design; this spec does not address it.
|
|
29
|
-
|
|
30
|
-
## Key Decisions (locked)
|
|
31
|
-
|
|
32
|
-
| Decision | Choice |
|
|
33
|
-
|---|---|
|
|
34
|
-
| Verify authority | The **implementer** runner's `success` flag is advisory; **verify decides**. Run verify regardless of runner exit, block only on verify failure. |
|
|
35
|
-
| Runner-fail + verify-pass | Proceed (commit) — the exit code was a ghost. Note it via the reporter / summary. |
|
|
36
|
-
| Runner-fail + verify-fail | Block, with a reason naming both signals. |
|
|
37
|
-
| Reviewer runner | **Unchanged** — its non-zero exit *is* the reject verdict (there is no separate verify for the review). |
|
|
38
|
-
| Flaky tolerance | `retryingVerifier(inner, retries)` re-runs a failing verify up to `retries` times; first pass wins. |
|
|
39
|
-
| Retry default | `config.verify.retries` default **1** (one retry). `0` = strict/no-retry. A real failure still fails (twice). |
|
|
40
|
-
| Out of scope | Cross-story regression repair; changing the reviewer's exit-code semantics. |
|
|
41
|
-
|
|
42
|
-
## Architecture
|
|
43
|
-
|
|
44
|
-
### 1. `src/loop/verify.ts` — `retryingVerifier`
|
|
45
|
-
```ts
|
|
46
|
-
export function retryingVerifier(inner: Verifier, retries: number): Verifier {
|
|
47
|
-
return (targetDir) => {
|
|
48
|
-
let last = inner(targetDir)
|
|
49
|
-
let attempt = 0
|
|
50
|
-
while (!last.passed && attempt < retries) {
|
|
51
|
-
attempt++
|
|
52
|
-
last = inner(targetDir)
|
|
53
|
-
}
|
|
54
|
-
if (last.passed && attempt > 0) {
|
|
55
|
-
return { passed: true, summary: `${last.summary} (passed on retry ${attempt})` }
|
|
56
|
-
}
|
|
57
|
-
return attempt > 0 && !last.passed
|
|
58
|
-
? { passed: false, summary: `${last.summary} (still failing after ${attempt} retr${attempt === 1 ? 'y' : 'ies'})` }
|
|
59
|
-
: last
|
|
60
|
-
}
|
|
61
|
-
}
|
|
62
|
-
```
|
|
63
|
-
Pure, injectable `inner` for tests (no real command needed).
|
|
64
|
-
|
|
65
|
-
### 2. `src/loop/loop.ts` — verify is the gate
|
|
66
|
-
In **both** the isolate and non-isolate paths, restructure the implementer step so verify always
|
|
67
|
-
runs:
|
|
68
|
-
|
|
69
|
-
- Run the implementer runner (keep `result` for its summary).
|
|
70
|
-
- `reporter.phase('verifying')`; run `opts.verify(dir)`.
|
|
71
|
-
- If verify **fails**: block. Reason = verify summary, and if the runner *also* reported failure,
|
|
72
|
-
prepend that (`runner reported failure (<summary>); verify also red: <verify summary>`).
|
|
73
|
-
- If verify **passes**: proceed to review/commit even if `result.success` was false. When the
|
|
74
|
-
runner had reported failure, the committed decision/summary notes
|
|
75
|
-
`(runner exited non-zero but verify is green)` so the ghost is auditable.
|
|
76
|
-
|
|
77
|
-
The reviewer step is untouched: `if (!reviewResult.success) → block` stays (exit = verdict).
|
|
78
|
-
|
|
79
|
-
The Baustein-E invariant (no `passes:true` without a commit; decision rollback on commit failure)
|
|
80
|
-
and the leftover-hint are preserved exactly — only the implementer-failure gate moves from
|
|
81
|
-
"before verify" to "verify decides".
|
|
82
|
-
|
|
83
|
-
### 3. `src/retrofit/config.ts` — `verify.retries`
|
|
84
|
-
Extend the `verify` config object: `verify: { command: string; retries?: number }`.
|
|
85
|
-
Backwards-compatible (optional).
|
|
86
|
-
|
|
87
|
-
### 4. `src/loop/run-command.ts` — wire the retry
|
|
88
|
-
When building the verifier from config, wrap it:
|
|
89
|
-
`verify = retryingVerifier(commandVerifier(command), config.verify?.retries ?? 1)`.
|
|
90
|
-
The injected-verifier test path (`opts.verify`) is unchanged.
|
|
91
|
-
|
|
92
|
-
### 5. README + loop-spec
|
|
93
|
-
Document: verify is the source of truth (a spurious agent exit code can't block a green story),
|
|
94
|
-
and `verify.retries` (default 1) for flaky suites. Update `canon/loop/loop-spec.md` step 5.
|
|
95
|
-
|
|
96
|
-
## Testing (subagent-driven TDD)
|
|
97
|
-
- **verify.ts:** `retryingVerifier` passes immediately when inner passes (no retry); passes on
|
|
98
|
-
retry N when inner fails then passes; fails after exhausting retries; `retries: 0` = single
|
|
99
|
-
shot; summary notes the retry/exhaustion. Inner is a stub Verifier.
|
|
100
|
-
- **loop.ts:** runner-fail + verify-pass → story is committed + `passes:true` (the SE5 case);
|
|
101
|
-
runner-fail + verify-fail → blocked with a combined reason; runner-pass + verify-fail →
|
|
102
|
-
blocked (unchanged); the reviewer still blocks on its own failure; existing invariant tests
|
|
103
|
-
(commit-failure revert, leftover hint) stay green; applies to both isolate and non-isolate.
|
|
104
|
-
- **config.ts:** `verify.retries` accepted + optional.
|
|
105
|
-
- **run-command.ts:** the built verifier is wrapped with the resolved retry count (assert via a
|
|
106
|
-
small seam or by config round-trip).
|
|
107
|
-
- Full suite green; `tsc` clean.
|
|
108
|
-
|
|
109
|
-
## What this would have done for the run
|
|
110
|
-
- **SE5:** runner exits 127 but suite is 769/769 green → verify passes → story auto-commits. No
|
|
111
|
-
manual finalize.
|
|
112
|
-
- **T4:** the load-flake fails once, the retry passes → story proceeds. No manual `vitest.config`
|
|
113
|
-
patch, no false block.
|
|
1
|
+
# Baustein H — Loop Robustness (verify-as-truth + flaky-tolerant verify)
|
|
2
|
+
|
|
3
|
+
**Status:** Design approved 2026-06-29
|
|
4
|
+
**Component:** Yoke (🐂)
|
|
5
|
+
**Relates to:** [[harness-build-progress]], [[harness-loop-technique]]
|
|
6
|
+
|
|
7
|
+
## Problem & Goal
|
|
8
|
+
|
|
9
|
+
Evidence from a real autonomous run (NewMarket, 52/57 stories): the loop blocked and needed
|
|
10
|
+
human intervention on a handful of stories. Two of those block causes are **harness defects**,
|
|
11
|
+
not project defects, and both are cheap to fix:
|
|
12
|
+
|
|
13
|
+
1. **Runner exit-code is treated as ground truth (SE5).** The story was implemented and the full
|
|
14
|
+
suite was green (769/769), but the `claude` `.cmd` wrapper on Windows exited **127**. The loop
|
|
15
|
+
checks `result.success` *before* running verify, so it **blocked a successful story** on a
|
|
16
|
+
spurious exit code and didn't auto-commit. This contradicts the loop's own stated philosophy
|
|
17
|
+
(Baustein C2): *"independent verification is the source of truth, not the agent's exit code."*
|
|
18
|
+
|
|
19
|
+
2. **No flaky-test tolerance (T4).** Under autonomous load, a heavy background task caused an
|
|
20
|
+
async test to time out — a **false** red, not a real defect. The single-shot verify gate
|
|
21
|
+
blocked the story. The user had to manually add `retry: 2` to the project's `vitest.config`.
|
|
22
|
+
That resilience belongs in the harness.
|
|
23
|
+
|
|
24
|
+
**Goal:** make the loop trust **verify**, not the agent's exit code, and tolerate transient
|
|
25
|
+
flakes — closing two real block causes without weakening the gate.
|
|
26
|
+
|
|
27
|
+
**Out of scope (deferred):** cross-story regression repair (S5/S6 — a later story breaking an
|
|
28
|
+
earlier story's tests). That needs its own design; this spec does not address it.
|
|
29
|
+
|
|
30
|
+
## Key Decisions (locked)
|
|
31
|
+
|
|
32
|
+
| Decision | Choice |
|
|
33
|
+
|---|---|
|
|
34
|
+
| Verify authority | The **implementer** runner's `success` flag is advisory; **verify decides**. Run verify regardless of runner exit, block only on verify failure. |
|
|
35
|
+
| Runner-fail + verify-pass | Proceed (commit) — the exit code was a ghost. Note it via the reporter / summary. |
|
|
36
|
+
| Runner-fail + verify-fail | Block, with a reason naming both signals. |
|
|
37
|
+
| Reviewer runner | **Unchanged** — its non-zero exit *is* the reject verdict (there is no separate verify for the review). |
|
|
38
|
+
| Flaky tolerance | `retryingVerifier(inner, retries)` re-runs a failing verify up to `retries` times; first pass wins. |
|
|
39
|
+
| Retry default | `config.verify.retries` default **1** (one retry). `0` = strict/no-retry. A real failure still fails (twice). |
|
|
40
|
+
| Out of scope | Cross-story regression repair; changing the reviewer's exit-code semantics. |
|
|
41
|
+
|
|
42
|
+
## Architecture
|
|
43
|
+
|
|
44
|
+
### 1. `src/loop/verify.ts` — `retryingVerifier`
|
|
45
|
+
```ts
|
|
46
|
+
export function retryingVerifier(inner: Verifier, retries: number): Verifier {
|
|
47
|
+
return (targetDir) => {
|
|
48
|
+
let last = inner(targetDir)
|
|
49
|
+
let attempt = 0
|
|
50
|
+
while (!last.passed && attempt < retries) {
|
|
51
|
+
attempt++
|
|
52
|
+
last = inner(targetDir)
|
|
53
|
+
}
|
|
54
|
+
if (last.passed && attempt > 0) {
|
|
55
|
+
return { passed: true, summary: `${last.summary} (passed on retry ${attempt})` }
|
|
56
|
+
}
|
|
57
|
+
return attempt > 0 && !last.passed
|
|
58
|
+
? { passed: false, summary: `${last.summary} (still failing after ${attempt} retr${attempt === 1 ? 'y' : 'ies'})` }
|
|
59
|
+
: last
|
|
60
|
+
}
|
|
61
|
+
}
|
|
62
|
+
```
|
|
63
|
+
Pure, injectable `inner` for tests (no real command needed).
|
|
64
|
+
|
|
65
|
+
### 2. `src/loop/loop.ts` — verify is the gate
|
|
66
|
+
In **both** the isolate and non-isolate paths, restructure the implementer step so verify always
|
|
67
|
+
runs:
|
|
68
|
+
|
|
69
|
+
- Run the implementer runner (keep `result` for its summary).
|
|
70
|
+
- `reporter.phase('verifying')`; run `opts.verify(dir)`.
|
|
71
|
+
- If verify **fails**: block. Reason = verify summary, and if the runner *also* reported failure,
|
|
72
|
+
prepend that (`runner reported failure (<summary>); verify also red: <verify summary>`).
|
|
73
|
+
- If verify **passes**: proceed to review/commit even if `result.success` was false. When the
|
|
74
|
+
runner had reported failure, the committed decision/summary notes
|
|
75
|
+
`(runner exited non-zero but verify is green)` so the ghost is auditable.
|
|
76
|
+
|
|
77
|
+
The reviewer step is untouched: `if (!reviewResult.success) → block` stays (exit = verdict).
|
|
78
|
+
|
|
79
|
+
The Baustein-E invariant (no `passes:true` without a commit; decision rollback on commit failure)
|
|
80
|
+
and the leftover-hint are preserved exactly — only the implementer-failure gate moves from
|
|
81
|
+
"before verify" to "verify decides".
|
|
82
|
+
|
|
83
|
+
### 3. `src/retrofit/config.ts` — `verify.retries`
|
|
84
|
+
Extend the `verify` config object: `verify: { command: string; retries?: number }`.
|
|
85
|
+
Backwards-compatible (optional).
|
|
86
|
+
|
|
87
|
+
### 4. `src/loop/run-command.ts` — wire the retry
|
|
88
|
+
When building the verifier from config, wrap it:
|
|
89
|
+
`verify = retryingVerifier(commandVerifier(command), config.verify?.retries ?? 1)`.
|
|
90
|
+
The injected-verifier test path (`opts.verify`) is unchanged.
|
|
91
|
+
|
|
92
|
+
### 5. README + loop-spec
|
|
93
|
+
Document: verify is the source of truth (a spurious agent exit code can't block a green story),
|
|
94
|
+
and `verify.retries` (default 1) for flaky suites. Update `canon/loop/loop-spec.md` step 5.
|
|
95
|
+
|
|
96
|
+
## Testing (subagent-driven TDD)
|
|
97
|
+
- **verify.ts:** `retryingVerifier` passes immediately when inner passes (no retry); passes on
|
|
98
|
+
retry N when inner fails then passes; fails after exhausting retries; `retries: 0` = single
|
|
99
|
+
shot; summary notes the retry/exhaustion. Inner is a stub Verifier.
|
|
100
|
+
- **loop.ts:** runner-fail + verify-pass → story is committed + `passes:true` (the SE5 case);
|
|
101
|
+
runner-fail + verify-fail → blocked with a combined reason; runner-pass + verify-fail →
|
|
102
|
+
blocked (unchanged); the reviewer still blocks on its own failure; existing invariant tests
|
|
103
|
+
(commit-failure revert, leftover hint) stay green; applies to both isolate and non-isolate.
|
|
104
|
+
- **config.ts:** `verify.retries` accepted + optional.
|
|
105
|
+
- **run-command.ts:** the built verifier is wrapped with the resolved retry count (assert via a
|
|
106
|
+
small seam or by config round-trip).
|
|
107
|
+
- Full suite green; `tsc` clean.
|
|
108
|
+
|
|
109
|
+
## What this would have done for the run
|
|
110
|
+
- **SE5:** runner exits 127 but suite is 769/769 green → verify passes → story auto-commits. No
|
|
111
|
+
manual finalize.
|
|
112
|
+
- **T4:** the load-flake fails once, the retry passes → story proceeds. No manual `vitest.config`
|
|
113
|
+
patch, no false block.
|
|
@@ -1,98 +1,98 @@
|
|
|
1
|
-
# Baustein I — Visual & Design Verification
|
|
2
|
-
|
|
3
|
-
**Status:** Design approved 2026-06-30 (autonomous)
|
|
4
|
-
**Component:** Yoke (🐂)
|
|
5
|
-
**Relates to:** [[harness-build-progress]], [[readme-always-update]]
|
|
6
|
-
**Inspired by:** [vibecoded-design-tells](https://github.com/JCarterJohnson/vibecoded-design-tells) (MIT © 2026 Carter Johnson) — a data-ranked study of the visual "tells" of AI-generated UIs. Yoke implements the *idea* natively in TypeScript and credits the research; no code or data is copied.
|
|
7
|
-
|
|
8
|
-
## Problem & Goal
|
|
9
|
-
|
|
10
|
-
Yoke's verify gate is **code-only** (`tsc` + unit/component tests). It does not check that the UI is **visually sound** (free of generic AI-slop design) or that **user flows actually work end-to-end**. Evidence: in the real NewMarket run, integration/visual bugs (unwired auth pages, a seed id-collision, an old "purple #6c5ce7 dark theme" — itself a classic AI-slop tell) slipped past the unit-test gate and were only caught by a later manual QA sweep.
|
|
11
|
-
|
|
12
|
-
**Goal:** add a **visual & design verification layer** that plugs into the existing verify model (so it's gated, not advisory) and is honest about cost:
|
|
13
|
-
1. **Mechanical:** a static **design-slop scanner** (`yoke design-scan`) that flags the high-signal AI-slop tells and gates on exit code.
|
|
14
|
-
2. **Methodology:** two canon skills — `unslop-ui` (the design rubric) and `visual-verification` (compose a verify pipeline: types + unit + design-scan + a Playwright flow-smoke; capture video only on failure).
|
|
15
|
-
|
|
16
|
-
Both integrate through one idea: **the project's `verify.command` becomes a pipeline**, and Baustein-H's verify-as-truth makes those gates authoritative.
|
|
17
|
-
|
|
18
|
-
## Key Decisions (locked)
|
|
19
|
-
|
|
20
|
-
| Decision | Choice |
|
|
21
|
-
|---|---|
|
|
22
|
-
| Design-slop detection | A **TS-native** scanner built into the `yoke` CLI (`yoke design-scan`), not a port of the upstream Python; high-precision static heuristics |
|
|
23
|
-
| Gate model | Exit non-zero when the weighted tell-score exceeds `--max` (default **4**); `--report` lists without failing |
|
|
24
|
-
| Flow / video | **Methodology, not CLI** — the agent drives the wired Playwright MCP per the `visual-verification` skill. Yoke gates + guides; it does not embed a browser. Video capture is **opt-in, on failure only** (token-aware). |
|
|
25
|
-
| Skills | `unslop-ui` (rubric) + `visual-verification` (pipeline + flow-smoke + video-on-failure) → all 3 agents |
|
|
26
|
-
| Attribution | Credit vibecoded-design-tells (MIT) in `ATTRIBUTION.md` + README + the skill |
|
|
27
|
-
| Canon count | 24 → **26** skills |
|
|
28
|
-
| Out of scope (YAGNI) | Embedding Playwright/a browser in the Yoke CLI; structural-layout detection ("centered hero + 3 cards") — left to the rubric + agent eye; auto-fixing slop (the agent fixes, guided by the skill) |
|
|
29
|
-
|
|
30
|
-
## Architecture
|
|
31
|
-
|
|
32
|
-
### 1. `src/scan/design.ts` (new) — the static scanner
|
|
33
|
-
Pure + unit-testable. Walks the project's source and scores AI-slop tells.
|
|
34
|
-
|
|
35
|
-
```ts
|
|
36
|
-
export interface Tell { name: string; weight: number; test: (line: string) => boolean; hint: string }
|
|
37
|
-
export interface Finding { file: string; line: number; tell: string; hint: string; text: string }
|
|
38
|
-
export interface ScanResult { findings: Finding[]; score: number }
|
|
39
|
-
|
|
40
|
-
export const TELLS: Tell[] // the curated tell set (below)
|
|
41
|
-
export function scanText(text: string, tells?: Tell[]): { line: number; tell: Tell; text: string }[]
|
|
42
|
-
export function scanDir(dir: string, tells?: Tell[]): ScanResult // walks files, aggregates
|
|
43
|
-
```
|
|
44
|
-
|
|
45
|
-
**Curated high-precision tells** (each match adds `weight` to the score):
|
|
46
|
-
|
|
47
|
-
| Tell | weight | matches (case-insensitive) | hint |
|
|
48
|
-
|---|---|---|---|
|
|
49
|
-
| `ai-purple` | 2 | hex `#6c5ce7\|#7c3aed\|#8b5cf6\|#a855f7\|#9333ea`, or Tailwind `(from\|via\|to)-(purple\|violet\|fuchsia)-(4\|5\|6\|7)00` | AI-purple is the #1 vibecoded tell — choose a real brand color |
|
|
50
|
-
| `gradient-clip-text` | 2 | a line containing both `bg-clip-text` and `text-transparent`, or CSS `-webkit-background-clip:\s*text` near a gradient | gradient hero text reads as AI-slop — solid color + weight instead |
|
|
51
|
-
| `neon-glow` | 2 | Tailwind `(shadow\|drop-shadow)-\[0_0_`, or CSS `box-shadow:[^;]*0\s+0\s+\d{2,}px` with a color | neon glow is a tell — use subtle, neutral elevation |
|
|
52
|
-
| `gradient-overload` | 1 | `bg-gradient-to-` or CSS `linear-gradient(` | gradients everywhere flatten hierarchy — use them sparingly |
|
|
53
|
-
| `emoji-icon` | 1 | an emoji (unicode pictographic) inside a `.tsx/.jsx` line that also contains `<button`, `<a `, `aria-hidden`, or a JSX `>…<` icon slot | emoji-as-icons is a tell — use a real icon set |
|
|
54
|
-
|
|
55
|
-
File walk: extensions `.css .scss .tsx .jsx .ts .js .html .vue .svelte .astro`; skip `node_modules`, `dist`, `.next`, `build`, `.yoke`, `coverage`, `.git`. The tell set is the default but injectable (for tests). Heuristic by design — documented as high-signal, not exhaustive.
|
|
56
|
-
|
|
57
|
-
### 2. `src/cli.ts` — `yoke design-scan [dir] [--max=N] [--report]`
|
|
58
|
-
- Runs `scanDir(dir)`, prints findings grouped by tell as `file:line <tell> — <hint>`.
|
|
59
|
-
- `--report`: print + summary, **always exit 0** (advisory).
|
|
60
|
-
- default (gate): print + summary; **exit 1 if `score > max`** (default `max=4`), else 0. So a couple incidental matches pass; pervasive slop fails. Designed to sit in a verify pipeline.
|
|
61
|
-
- A small body extracted to `runDesignScan(dir, { max, report }): number` for testability.
|
|
62
|
-
|
|
63
|
-
### 3. `canon/skills/unslop-ui/SKILL.md` (new) — the rubric
|
|
64
|
-
Agent-facing. Lists the ranked tells (AI-purple gradients, gradient hero text, neon glow, emoji-as-icons, shadcn defaults left unchanged, "centered hero + three cards", homogeneous spacing) and how to fix each. Instructs: before finishing UI work, run `yoke design-scan .` and resolve findings; also apply the structural items the scanner can't see. Credits the research.
|
|
65
|
-
|
|
66
|
-
### 4. `canon/skills/visual-verification/SKILL.md` (new) — ties it together
|
|
67
|
-
Methodology for UI projects:
|
|
68
|
-
- **Compose the verify pipeline** so the loop gate covers more than units: `verify.command` chains types → unit tests → `yoke design-scan .` → a Playwright flow-smoke. (Baustein-H makes these authoritative.)
|
|
69
|
-
- **Flow-smoke via the wired Playwright MCP:** load the key routes against the dev server, assert they render and the console has **no errors**, screenshot each. This catches the "unwired page / runtime crash" bugs unit tests miss.
|
|
70
|
-
- **Video only when necessary:** capture a video of a flow **only on failure** (or when explicitly debugging a UX issue), then analyse it — keeps tokens down. Never record every run.
|
|
71
|
-
|
|
72
|
-
### 5. Wiring
|
|
73
|
-
- `canon/manifest.yaml`: add `unslop-ui` + `visual-verification` (kind: methodology).
|
|
74
|
-
- `canon/skills/ATTRIBUTION.md`: credit vibecoded-design-tells (MIT © Carter Johnson).
|
|
75
|
-
- `README.md` (**mandatory**): a "Visual & design verification" section (the scanner + the two skills + the pipeline idea), the catalog updated (24 → 26, methodology group +2), and the test-count badge synced.
|
|
76
|
-
|
|
77
|
-
## Data flow (gate)
|
|
78
|
-
```
|
|
79
|
-
verify.command = tsc --noEmit && vitest run && yoke design-scan . && <playwright flow-smoke>
|
|
80
|
-
│ exit 1 if slop-score > max
|
|
81
|
-
loop verify (Baustein H: verify is the source of truth) ──► block on any red step
|
|
82
|
-
```
|
|
83
|
-
|
|
84
|
-
## Testing (subagent-driven TDD)
|
|
85
|
-
- **design.ts:** `scanText` flags each tell with correct line + weight; clean text → no findings; `scanDir` walks + skips ignored dirs + aggregates score; injected tell-set works; emoji/purple/clip/glow/gradient cases each covered; a known-clean snippet scores 0.
|
|
86
|
-
- **cli:** `runDesignScan` exits 1 when score > max, 0 when ≤ max, 0 always in `--report`; `--max` parsed; bad `--max` rejected.
|
|
87
|
-
- **canon:** `unslop-ui` + `visual-verification` registered; `validateCanon` stays zero-error; real-canon asserts both present.
|
|
88
|
-
- Full suite green; `tsc` clean.
|
|
89
|
-
|
|
90
|
-
## What this would have caught
|
|
91
|
-
- NewMarket's old **purple `#6c5ce7` dark theme** → `ai-purple` tell, scored, flagged before it shipped.
|
|
92
|
-
- **Unwired auth pages / runtime crashes** → the flow-smoke (render + no console errors) gate, not the unit tests.
|
|
93
|
-
|
|
94
|
-
## Non-goals (YAGNI)
|
|
95
|
-
- No browser embedded in the Yoke CLI (Playwright MCP is the agent's tool).
|
|
96
|
-
- No always-on video (opt-in, on failure).
|
|
97
|
-
- No auto-rewrite of slop (the agent fixes via the rubric).
|
|
98
|
-
- No structural-layout static detection (rubric + agent eye).
|
|
1
|
+
# Baustein I — Visual & Design Verification
|
|
2
|
+
|
|
3
|
+
**Status:** Design approved 2026-06-30 (autonomous)
|
|
4
|
+
**Component:** Yoke (🐂)
|
|
5
|
+
**Relates to:** [[harness-build-progress]], [[readme-always-update]]
|
|
6
|
+
**Inspired by:** [vibecoded-design-tells](https://github.com/JCarterJohnson/vibecoded-design-tells) (MIT © 2026 Carter Johnson) — a data-ranked study of the visual "tells" of AI-generated UIs. Yoke implements the *idea* natively in TypeScript and credits the research; no code or data is copied.
|
|
7
|
+
|
|
8
|
+
## Problem & Goal
|
|
9
|
+
|
|
10
|
+
Yoke's verify gate is **code-only** (`tsc` + unit/component tests). It does not check that the UI is **visually sound** (free of generic AI-slop design) or that **user flows actually work end-to-end**. Evidence: in the real NewMarket run, integration/visual bugs (unwired auth pages, a seed id-collision, an old "purple #6c5ce7 dark theme" — itself a classic AI-slop tell) slipped past the unit-test gate and were only caught by a later manual QA sweep.
|
|
11
|
+
|
|
12
|
+
**Goal:** add a **visual & design verification layer** that plugs into the existing verify model (so it's gated, not advisory) and is honest about cost:
|
|
13
|
+
1. **Mechanical:** a static **design-slop scanner** (`yoke design-scan`) that flags the high-signal AI-slop tells and gates on exit code.
|
|
14
|
+
2. **Methodology:** two canon skills — `unslop-ui` (the design rubric) and `visual-verification` (compose a verify pipeline: types + unit + design-scan + a Playwright flow-smoke; capture video only on failure).
|
|
15
|
+
|
|
16
|
+
Both integrate through one idea: **the project's `verify.command` becomes a pipeline**, and Baustein-H's verify-as-truth makes those gates authoritative.
|
|
17
|
+
|
|
18
|
+
## Key Decisions (locked)
|
|
19
|
+
|
|
20
|
+
| Decision | Choice |
|
|
21
|
+
|---|---|
|
|
22
|
+
| Design-slop detection | A **TS-native** scanner built into the `yoke` CLI (`yoke design-scan`), not a port of the upstream Python; high-precision static heuristics |
|
|
23
|
+
| Gate model | Exit non-zero when the weighted tell-score exceeds `--max` (default **4**); `--report` lists without failing |
|
|
24
|
+
| Flow / video | **Methodology, not CLI** — the agent drives the wired Playwright MCP per the `visual-verification` skill. Yoke gates + guides; it does not embed a browser. Video capture is **opt-in, on failure only** (token-aware). |
|
|
25
|
+
| Skills | `unslop-ui` (rubric) + `visual-verification` (pipeline + flow-smoke + video-on-failure) → all 3 agents |
|
|
26
|
+
| Attribution | Credit vibecoded-design-tells (MIT) in `ATTRIBUTION.md` + README + the skill |
|
|
27
|
+
| Canon count | 24 → **26** skills |
|
|
28
|
+
| Out of scope (YAGNI) | Embedding Playwright/a browser in the Yoke CLI; structural-layout detection ("centered hero + 3 cards") — left to the rubric + agent eye; auto-fixing slop (the agent fixes, guided by the skill) |
|
|
29
|
+
|
|
30
|
+
## Architecture
|
|
31
|
+
|
|
32
|
+
### 1. `src/scan/design.ts` (new) — the static scanner
|
|
33
|
+
Pure + unit-testable. Walks the project's source and scores AI-slop tells.
|
|
34
|
+
|
|
35
|
+
```ts
|
|
36
|
+
export interface Tell { name: string; weight: number; test: (line: string) => boolean; hint: string }
|
|
37
|
+
export interface Finding { file: string; line: number; tell: string; hint: string; text: string }
|
|
38
|
+
export interface ScanResult { findings: Finding[]; score: number }
|
|
39
|
+
|
|
40
|
+
export const TELLS: Tell[] // the curated tell set (below)
|
|
41
|
+
export function scanText(text: string, tells?: Tell[]): { line: number; tell: Tell; text: string }[]
|
|
42
|
+
export function scanDir(dir: string, tells?: Tell[]): ScanResult // walks files, aggregates
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
**Curated high-precision tells** (each match adds `weight` to the score):
|
|
46
|
+
|
|
47
|
+
| Tell | weight | matches (case-insensitive) | hint |
|
|
48
|
+
|---|---|---|---|
|
|
49
|
+
| `ai-purple` | 2 | hex `#6c5ce7\|#7c3aed\|#8b5cf6\|#a855f7\|#9333ea`, or Tailwind `(from\|via\|to)-(purple\|violet\|fuchsia)-(4\|5\|6\|7)00` | AI-purple is the #1 vibecoded tell — choose a real brand color |
|
|
50
|
+
| `gradient-clip-text` | 2 | a line containing both `bg-clip-text` and `text-transparent`, or CSS `-webkit-background-clip:\s*text` near a gradient | gradient hero text reads as AI-slop — solid color + weight instead |
|
|
51
|
+
| `neon-glow` | 2 | Tailwind `(shadow\|drop-shadow)-\[0_0_`, or CSS `box-shadow:[^;]*0\s+0\s+\d{2,}px` with a color | neon glow is a tell — use subtle, neutral elevation |
|
|
52
|
+
| `gradient-overload` | 1 | `bg-gradient-to-` or CSS `linear-gradient(` | gradients everywhere flatten hierarchy — use them sparingly |
|
|
53
|
+
| `emoji-icon` | 1 | an emoji (unicode pictographic) inside a `.tsx/.jsx` line that also contains `<button`, `<a `, `aria-hidden`, or a JSX `>…<` icon slot | emoji-as-icons is a tell — use a real icon set |
|
|
54
|
+
|
|
55
|
+
File walk: extensions `.css .scss .tsx .jsx .ts .js .html .vue .svelte .astro`; skip `node_modules`, `dist`, `.next`, `build`, `.yoke`, `coverage`, `.git`. The tell set is the default but injectable (for tests). Heuristic by design — documented as high-signal, not exhaustive.
|
|
56
|
+
|
|
57
|
+
### 2. `src/cli.ts` — `yoke design-scan [dir] [--max=N] [--report]`
|
|
58
|
+
- Runs `scanDir(dir)`, prints findings grouped by tell as `file:line <tell> — <hint>`.
|
|
59
|
+
- `--report`: print + summary, **always exit 0** (advisory).
|
|
60
|
+
- default (gate): print + summary; **exit 1 if `score > max`** (default `max=4`), else 0. So a couple incidental matches pass; pervasive slop fails. Designed to sit in a verify pipeline.
|
|
61
|
+
- A small body extracted to `runDesignScan(dir, { max, report }): number` for testability.
|
|
62
|
+
|
|
63
|
+
### 3. `canon/skills/unslop-ui/SKILL.md` (new) — the rubric
|
|
64
|
+
Agent-facing. Lists the ranked tells (AI-purple gradients, gradient hero text, neon glow, emoji-as-icons, shadcn defaults left unchanged, "centered hero + three cards", homogeneous spacing) and how to fix each. Instructs: before finishing UI work, run `yoke design-scan .` and resolve findings; also apply the structural items the scanner can't see. Credits the research.
|
|
65
|
+
|
|
66
|
+
### 4. `canon/skills/visual-verification/SKILL.md` (new) — ties it together
|
|
67
|
+
Methodology for UI projects:
|
|
68
|
+
- **Compose the verify pipeline** so the loop gate covers more than units: `verify.command` chains types → unit tests → `yoke design-scan .` → a Playwright flow-smoke. (Baustein-H makes these authoritative.)
|
|
69
|
+
- **Flow-smoke via the wired Playwright MCP:** load the key routes against the dev server, assert they render and the console has **no errors**, screenshot each. This catches the "unwired page / runtime crash" bugs unit tests miss.
|
|
70
|
+
- **Video only when necessary:** capture a video of a flow **only on failure** (or when explicitly debugging a UX issue), then analyse it — keeps tokens down. Never record every run.
|
|
71
|
+
|
|
72
|
+
### 5. Wiring
|
|
73
|
+
- `canon/manifest.yaml`: add `unslop-ui` + `visual-verification` (kind: methodology).
|
|
74
|
+
- `canon/skills/ATTRIBUTION.md`: credit vibecoded-design-tells (MIT © Carter Johnson).
|
|
75
|
+
- `README.md` (**mandatory**): a "Visual & design verification" section (the scanner + the two skills + the pipeline idea), the catalog updated (24 → 26, methodology group +2), and the test-count badge synced.
|
|
76
|
+
|
|
77
|
+
## Data flow (gate)
|
|
78
|
+
```
|
|
79
|
+
verify.command = tsc --noEmit && vitest run && yoke design-scan . && <playwright flow-smoke>
|
|
80
|
+
│ exit 1 if slop-score > max
|
|
81
|
+
loop verify (Baustein H: verify is the source of truth) ──► block on any red step
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
## Testing (subagent-driven TDD)
|
|
85
|
+
- **design.ts:** `scanText` flags each tell with correct line + weight; clean text → no findings; `scanDir` walks + skips ignored dirs + aggregates score; injected tell-set works; emoji/purple/clip/glow/gradient cases each covered; a known-clean snippet scores 0.
|
|
86
|
+
- **cli:** `runDesignScan` exits 1 when score > max, 0 when ≤ max, 0 always in `--report`; `--max` parsed; bad `--max` rejected.
|
|
87
|
+
- **canon:** `unslop-ui` + `visual-verification` registered; `validateCanon` stays zero-error; real-canon asserts both present.
|
|
88
|
+
- Full suite green; `tsc` clean.
|
|
89
|
+
|
|
90
|
+
## What this would have caught
|
|
91
|
+
- NewMarket's old **purple `#6c5ce7` dark theme** → `ai-purple` tell, scored, flagged before it shipped.
|
|
92
|
+
- **Unwired auth pages / runtime crashes** → the flow-smoke (render + no console errors) gate, not the unit tests.
|
|
93
|
+
|
|
94
|
+
## Non-goals (YAGNI)
|
|
95
|
+
- No browser embedded in the Yoke CLI (Playwright MCP is the agent's tool).
|
|
96
|
+
- No always-on video (opt-in, on failure).
|
|
97
|
+
- No auto-rewrite of slop (the agent fixes via the rubric).
|
|
98
|
+
- No structural-layout static detection (rubric + agent eye).
|