@hecer/yoke 1.11.0 → 1.13.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +13 -13
- package/.codex-plugin/plugin.json +7 -7
- package/CHANGELOG.md +435 -398
- package/README.md +943 -915
- package/TODOS.md +5 -5
- package/agents/docs.toml +6 -6
- package/agents/implementer.toml +6 -6
- package/agents/reviewer.toml +6 -6
- package/agents/security.toml +6 -6
- package/bench/README.md +86 -86
- package/bench/RESULTS.md +35 -35
- package/bench/output-compaction.mjs +65 -65
- package/bench/result-schema.mjs +12 -12
- package/bench/results/claude-2026-07-27T18-03-26.json +50 -50
- package/bench/results/codex-unavailable-1785175418318.json +15 -15
- package/bench/results/gemini-2026-07-27T18-03-44.json +46 -46
- package/bench/run-matrix.mjs +26 -26
- package/bench/run.mjs +106 -106
- package/canon/AGENTS.md +30 -30
- package/canon/context/DECISIONS.md +4 -4
- package/canon/context/GLOSSARY.md +11 -11
- package/canon/context/KNOWLEDGE.md +4 -4
- package/canon/context/PROJECT.md +15 -15
- package/canon/loop/loop-spec.md +65 -65
- package/canon/loop/prd.schema.md +41 -41
- package/canon/manifest.yaml +59 -59
- package/canon/policy/gates.md +7 -7
- package/canon/policy/roles.md +9 -9
- package/canon/skills/ATTRIBUTION.md +99 -99
- package/canon/skills/authoring-prd/SKILL.md +57 -57
- package/canon/skills/brainstorming/SKILL.md +164 -164
- package/canon/skills/codebase-design/DEEPENING.md +15 -15
- package/canon/skills/codebase-design/DESIGN-IT-TWICE.md +12 -12
- package/canon/skills/codebase-design/SKILL.md +39 -39
- package/canon/skills/dispatching-parallel-agents/SKILL.md +182 -182
- package/canon/skills/document-release/SKILL.md +302 -302
- package/canon/skills/domain-modeling/ADR-FORMAT.md +19 -19
- package/canon/skills/domain-modeling/CONTEXT-FORMAT.md +39 -39
- package/canon/skills/domain-modeling/SKILL.md +35 -35
- package/canon/skills/executing-plans/SKILL.md +70 -70
- package/canon/skills/finishing-a-development-branch/SKILL.md +200 -200
- package/canon/skills/health/SKILL.md +177 -177
- package/canon/skills/maintaining-context/SKILL.md +34 -34
- package/canon/skills/minimal-code/SKILL.md +21 -21
- package/canon/skills/no-ai-slop/SKILL.md +103 -103
- package/canon/skills/no-ai-slop/eval.md +43 -43
- package/canon/skills/plan-ceo-review/SKILL.md +541 -541
- package/canon/skills/plan-eng-review/SKILL.md +362 -362
- package/canon/skills/receiving-code-review/SKILL.md +213 -213
- package/canon/skills/requesting-code-review/SKILL.md +105 -105
- package/canon/skills/resolving-merge-conflicts/SKILL.md +18 -18
- package/canon/skills/retro/SKILL.md +397 -397
- package/canon/skills/review/SKILL.md +246 -246
- package/canon/skills/ship/SKILL.md +691 -691
- package/canon/skills/subagent-driven-development/SKILL.md +277 -277
- package/canon/skills/systematic-debugging/SKILL.md +296 -296
- package/canon/skills/tdd/SKILL.md +371 -371
- package/canon/skills/unslop-ui/SKILL.md +34 -34
- package/canon/skills/using-git-worktrees/SKILL.md +218 -218
- package/canon/skills/verification-before-completion/SKILL.md +139 -139
- package/canon/skills/visual-verification/SKILL.md +54 -54
- package/canon/skills/workflow/SKILL.md +22 -22
- package/canon/skills/writing-for-agents/SKILL-MECHANICS.md +27 -27
- package/canon/skills/writing-for-agents/SKILL.md +42 -42
- package/canon/skills/writing-plans/SKILL.md +152 -152
- package/canon/skills/writing-skills/SKILL.md +655 -655
- package/canon/skills/yoke-retrofit/SKILL.md +26 -26
- package/canon/skills/yoke-workflow/SKILL.md +20 -20
- package/canon/tools/codex-rtk-hook.mjs +35 -35
- package/canon/tools/gemini-rtk-hook.mjs +25 -25
- package/canon/tools/graphify.md +3 -3
- package/canon/tools/playwright-mcp.md +3 -3
- package/canon/tools/qwen-rtk-hook.mjs +25 -0
- package/canon/tools/rtk.md +7 -7
- package/canon/tools/serena.md +6 -6
- package/dist/agents/catalog.js +7 -0
- package/dist/agents/contracts.js +3 -1
- package/dist/agents/host.js +5 -1
- package/dist/agents/process-streams.js +62 -0
- package/dist/agents/process.js +43 -3
- package/dist/agents/providers.js +61 -6
- package/dist/agents/telemetry.js +133 -37
- package/dist/canon/manifest.js +2 -1
- package/dist/change/inbox.js +1 -1
- package/dist/cli.js +30 -24
- package/dist/dashboard/page.js +122 -122
- package/dist/dashboard/panels.js +91 -91
- package/dist/goals/command.js +3 -2
- package/dist/loop/claims.js +2 -1
- package/dist/loop/decision.js +3 -2
- package/dist/loop/parallel-command.js +4 -2
- package/dist/loop/prd.js +2 -1
- package/dist/loop/reporter.js +1 -0
- package/dist/loop/run-command.js +31 -10
- package/dist/prd/command.js +19 -19
- package/dist/quality/candidate-comparison.js +6 -1
- package/dist/quality/command.js +17 -2
- package/dist/quality/types.js +6 -1
- package/dist/retrofit/apply.js +95 -2
- package/dist/retrofit/config.js +9 -1
- package/dist/retrofit/detect.js +8 -0
- package/dist/retrofit/plan.js +6 -0
- package/dist/retrofit/planners/claude.js +14 -14
- package/dist/retrofit/planners/kilo.js +44 -0
- package/dist/retrofit/planners/opencode.js +44 -0
- package/dist/retrofit/planners/pi.js +24 -0
- package/dist/retrofit/planners/qwen.js +3 -3
- package/dist/retrofit/preserve.js +2 -2
- package/dist/retrofit/qwen-settings.js +17 -0
- package/dist/retrofit/skill-actions.js +4 -1
- package/dist/retrofit/tools.js +8 -0
- package/dist/review/command.js +3 -2
- package/dist/review/verdict.js +1 -1
- package/dist/routing/capability.js +2 -2
- package/dist/routing/planning.js +2 -0
- package/dist/routing/registry.js +3 -1
- package/dist/routing/router.js +7 -3
- package/dist/setup/command.js +35 -11
- package/dist/setup/model-presets.js +48 -0
- package/docs/CAPABILITY-ROUTING.md +51 -51
- package/docs/DASHBOARD-EVOLUTION.md +33 -33
- package/docs/HARNESSES.md +81 -0
- package/docs/MIGRATING-TO-1.0.md +33 -33
- package/docs/MIGRATING-TO-1.1.md +27 -27
- package/docs/MIGRATING-TO-1.4.md +70 -70
- package/docs/PRODUCT-DIRECTION-2026-09-05.md +210 -210
- package/docs/PUBLISHING.md +114 -114
- package/docs/QWEN-MODEL-SUPPORT.md +142 -0
- package/docs/VERIFIED-PROJECTS-VALIDATION.md +29 -29
- package/docs/VERIFIED-PROJECTS.md +167 -167
- package/docs/superpowers/plans/2026-06-28-baustein-e-context-layer.md +981 -981
- package/docs/superpowers/plans/2026-06-29-baustein-f-routing.md +258 -258
- package/docs/superpowers/plans/2026-06-29-baustein-g-loop-observability.md +1006 -1006
- package/docs/superpowers/plans/2026-06-29-baustein-h-loop-robustness.md +374 -374
- package/docs/superpowers/plans/2026-06-30-baustein-i-visual-design-verification.md +450 -450
- package/docs/superpowers/plans/2026-07-02-baustein-k-zero-to-100-bootstrap.md +1024 -1024
- package/docs/superpowers/plans/2026-07-02-baustein-m-flow-smoke-proofs.md +574 -574
- package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -537
- package/docs/superpowers/plans/2026-08-16-artifact-backed-output-compaction.md +329 -329
- package/docs/superpowers/plans/2026-09-05-verified-projects.md +83 -83
- package/docs/superpowers/specs/2026-06-28-baustein-e-context-layer-design.md +146 -146
- package/docs/superpowers/specs/2026-06-29-baustein-f-routing-design.md +106 -106
- package/docs/superpowers/specs/2026-06-29-baustein-g-loop-observability-design.md +186 -186
- package/docs/superpowers/specs/2026-06-29-baustein-h-loop-robustness-design.md +113 -113
- package/docs/superpowers/specs/2026-06-30-baustein-i-visual-design-verification-design.md +98 -98
- package/docs/superpowers/specs/2026-07-02-baustein-k-zero-to-100-bootstrap-design.md +200 -200
- package/docs/superpowers/specs/2026-07-02-baustein-m-flow-smoke-proofs-design.md +155 -155
- package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -422
- package/docs/superpowers/specs/2026-08-16-artifact-backed-output-compaction-design.md +166 -166
- package/gemini-extension.json +6 -6
- package/hooks/hooks.json +19 -19
- package/package.json +91 -87
- package/dist/dashboard/discovery.js +0 -73
- package/docs/community-outreach-2026-08-20.md +0 -85
- package/docs/launch-copy-2026-08-21.md +0 -193
|
@@ -1,139 +1,139 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: verification-before-completion
|
|
3
|
-
description: Use when about to claim work is complete, fixed, or passing, before committing or creating PRs - requires running verification commands and confirming output before making any success claims; evidence before assertions always
|
|
4
|
-
---
|
|
5
|
-
|
|
6
|
-
# Verification Before Completion
|
|
7
|
-
|
|
8
|
-
## Overview
|
|
9
|
-
|
|
10
|
-
Claiming work is complete without verification is dishonesty, not efficiency.
|
|
11
|
-
|
|
12
|
-
**Core principle:** Evidence before claims, always.
|
|
13
|
-
|
|
14
|
-
**Violating the letter of this rule is violating the spirit of this rule.**
|
|
15
|
-
|
|
16
|
-
## The Iron Law
|
|
17
|
-
|
|
18
|
-
```
|
|
19
|
-
NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE
|
|
20
|
-
```
|
|
21
|
-
|
|
22
|
-
If you haven't run the verification command in this message, you cannot claim it passes.
|
|
23
|
-
|
|
24
|
-
## The Gate Function
|
|
25
|
-
|
|
26
|
-
```
|
|
27
|
-
BEFORE claiming any status or expressing satisfaction:
|
|
28
|
-
|
|
29
|
-
1. IDENTIFY: What command proves this claim?
|
|
30
|
-
2. RUN: Execute the FULL command (fresh, complete)
|
|
31
|
-
3. READ: Full output, check exit code, count failures
|
|
32
|
-
4. VERIFY: Does output confirm the claim?
|
|
33
|
-
- If NO: State actual status with evidence
|
|
34
|
-
- If YES: State claim WITH evidence
|
|
35
|
-
5. ONLY THEN: Make the claim
|
|
36
|
-
|
|
37
|
-
Skip any step = lying, not verifying
|
|
38
|
-
```
|
|
39
|
-
|
|
40
|
-
## Common Failures
|
|
41
|
-
|
|
42
|
-
| Claim | Requires | Not Sufficient |
|
|
43
|
-
|-------|----------|----------------|
|
|
44
|
-
| Tests pass | Test command output: 0 failures | Previous run, "should pass" |
|
|
45
|
-
| Linter clean | Linter output: 0 errors | Partial check, extrapolation |
|
|
46
|
-
| Build succeeds | Build command: exit 0 | Linter passing, logs look good |
|
|
47
|
-
| Bug fixed | Test original symptom: passes | Code changed, assumed fixed |
|
|
48
|
-
| Regression test works | Red-green cycle verified | Test passes once |
|
|
49
|
-
| Agent completed | VCS diff shows changes | Agent reports "success" |
|
|
50
|
-
| Requirements met | Line-by-line checklist | Tests passing |
|
|
51
|
-
|
|
52
|
-
## Red Flags - STOP
|
|
53
|
-
|
|
54
|
-
- Using "should", "probably", "seems to"
|
|
55
|
-
- Expressing satisfaction before verification ("Great!", "Perfect!", "Done!", etc.)
|
|
56
|
-
- About to commit/push/PR without verification
|
|
57
|
-
- Trusting agent success reports
|
|
58
|
-
- Relying on partial verification
|
|
59
|
-
- Thinking "just this once"
|
|
60
|
-
- Tired and wanting work over
|
|
61
|
-
- **ANY wording implying success without having run verification**
|
|
62
|
-
|
|
63
|
-
## Rationalization Prevention
|
|
64
|
-
|
|
65
|
-
| Excuse | Reality |
|
|
66
|
-
|--------|---------|
|
|
67
|
-
| "Should work now" | RUN the verification |
|
|
68
|
-
| "I'm confident" | Confidence ≠ evidence |
|
|
69
|
-
| "Just this once" | No exceptions |
|
|
70
|
-
| "Linter passed" | Linter ≠ compiler |
|
|
71
|
-
| "Agent said success" | Verify independently |
|
|
72
|
-
| "I'm tired" | Exhaustion ≠ excuse |
|
|
73
|
-
| "Partial check is enough" | Partial proves nothing |
|
|
74
|
-
| "Different words so rule doesn't apply" | Spirit over letter |
|
|
75
|
-
|
|
76
|
-
## Key Patterns
|
|
77
|
-
|
|
78
|
-
**Tests:**
|
|
79
|
-
```
|
|
80
|
-
✅ [Run test command] [See: 34/34 pass] "All tests pass"
|
|
81
|
-
❌ "Should pass now" / "Looks correct"
|
|
82
|
-
```
|
|
83
|
-
|
|
84
|
-
**Regression tests (TDD Red-Green):**
|
|
85
|
-
```
|
|
86
|
-
✅ Write → Run (pass) → Revert fix → Run (MUST FAIL) → Restore → Run (pass)
|
|
87
|
-
❌ "I've written a regression test" (without red-green verification)
|
|
88
|
-
```
|
|
89
|
-
|
|
90
|
-
**Build:**
|
|
91
|
-
```
|
|
92
|
-
✅ [Run build] [See: exit 0] "Build passes"
|
|
93
|
-
❌ "Linter passed" (linter doesn't check compilation)
|
|
94
|
-
```
|
|
95
|
-
|
|
96
|
-
**Requirements:**
|
|
97
|
-
```
|
|
98
|
-
✅ Re-read plan → Create checklist → Verify each → Report gaps or completion
|
|
99
|
-
❌ "Tests pass, phase complete"
|
|
100
|
-
```
|
|
101
|
-
|
|
102
|
-
**Agent delegation:**
|
|
103
|
-
```
|
|
104
|
-
✅ Agent reports success → Check VCS diff → Verify changes → Report actual state
|
|
105
|
-
❌ Trust agent report
|
|
106
|
-
```
|
|
107
|
-
|
|
108
|
-
## Why This Matters
|
|
109
|
-
|
|
110
|
-
From 24 failure memories:
|
|
111
|
-
- your human partner said "I don't believe you" - trust broken
|
|
112
|
-
- Undefined functions shipped - would crash
|
|
113
|
-
- Missing requirements shipped - incomplete features
|
|
114
|
-
- Time wasted on false completion → redirect → rework
|
|
115
|
-
- Violates: "Honesty is a core value. If you lie, you'll be replaced."
|
|
116
|
-
|
|
117
|
-
## When To Apply
|
|
118
|
-
|
|
119
|
-
**ALWAYS before:**
|
|
120
|
-
- ANY variation of success/completion claims
|
|
121
|
-
- ANY expression of satisfaction
|
|
122
|
-
- ANY positive statement about work state
|
|
123
|
-
- Committing, PR creation, task completion
|
|
124
|
-
- Moving to next task
|
|
125
|
-
- Delegating to agents
|
|
126
|
-
|
|
127
|
-
**Rule applies to:**
|
|
128
|
-
- Exact phrases
|
|
129
|
-
- Paraphrases and synonyms
|
|
130
|
-
- Implications of success
|
|
131
|
-
- ANY communication suggesting completion/correctness
|
|
132
|
-
|
|
133
|
-
## The Bottom Line
|
|
134
|
-
|
|
135
|
-
**No shortcuts for verification.**
|
|
136
|
-
|
|
137
|
-
Run the command. Read the output. THEN claim the result.
|
|
138
|
-
|
|
139
|
-
This is non-negotiable.
|
|
1
|
+
---
|
|
2
|
+
name: verification-before-completion
|
|
3
|
+
description: Use when about to claim work is complete, fixed, or passing, before committing or creating PRs - requires running verification commands and confirming output before making any success claims; evidence before assertions always
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Verification Before Completion
|
|
7
|
+
|
|
8
|
+
## Overview
|
|
9
|
+
|
|
10
|
+
Claiming work is complete without verification is dishonesty, not efficiency.
|
|
11
|
+
|
|
12
|
+
**Core principle:** Evidence before claims, always.
|
|
13
|
+
|
|
14
|
+
**Violating the letter of this rule is violating the spirit of this rule.**
|
|
15
|
+
|
|
16
|
+
## The Iron Law
|
|
17
|
+
|
|
18
|
+
```
|
|
19
|
+
NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
If you haven't run the verification command in this message, you cannot claim it passes.
|
|
23
|
+
|
|
24
|
+
## The Gate Function
|
|
25
|
+
|
|
26
|
+
```
|
|
27
|
+
BEFORE claiming any status or expressing satisfaction:
|
|
28
|
+
|
|
29
|
+
1. IDENTIFY: What command proves this claim?
|
|
30
|
+
2. RUN: Execute the FULL command (fresh, complete)
|
|
31
|
+
3. READ: Full output, check exit code, count failures
|
|
32
|
+
4. VERIFY: Does output confirm the claim?
|
|
33
|
+
- If NO: State actual status with evidence
|
|
34
|
+
- If YES: State claim WITH evidence
|
|
35
|
+
5. ONLY THEN: Make the claim
|
|
36
|
+
|
|
37
|
+
Skip any step = lying, not verifying
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
## Common Failures
|
|
41
|
+
|
|
42
|
+
| Claim | Requires | Not Sufficient |
|
|
43
|
+
|-------|----------|----------------|
|
|
44
|
+
| Tests pass | Test command output: 0 failures | Previous run, "should pass" |
|
|
45
|
+
| Linter clean | Linter output: 0 errors | Partial check, extrapolation |
|
|
46
|
+
| Build succeeds | Build command: exit 0 | Linter passing, logs look good |
|
|
47
|
+
| Bug fixed | Test original symptom: passes | Code changed, assumed fixed |
|
|
48
|
+
| Regression test works | Red-green cycle verified | Test passes once |
|
|
49
|
+
| Agent completed | VCS diff shows changes | Agent reports "success" |
|
|
50
|
+
| Requirements met | Line-by-line checklist | Tests passing |
|
|
51
|
+
|
|
52
|
+
## Red Flags - STOP
|
|
53
|
+
|
|
54
|
+
- Using "should", "probably", "seems to"
|
|
55
|
+
- Expressing satisfaction before verification ("Great!", "Perfect!", "Done!", etc.)
|
|
56
|
+
- About to commit/push/PR without verification
|
|
57
|
+
- Trusting agent success reports
|
|
58
|
+
- Relying on partial verification
|
|
59
|
+
- Thinking "just this once"
|
|
60
|
+
- Tired and wanting work over
|
|
61
|
+
- **ANY wording implying success without having run verification**
|
|
62
|
+
|
|
63
|
+
## Rationalization Prevention
|
|
64
|
+
|
|
65
|
+
| Excuse | Reality |
|
|
66
|
+
|--------|---------|
|
|
67
|
+
| "Should work now" | RUN the verification |
|
|
68
|
+
| "I'm confident" | Confidence ≠ evidence |
|
|
69
|
+
| "Just this once" | No exceptions |
|
|
70
|
+
| "Linter passed" | Linter ≠ compiler |
|
|
71
|
+
| "Agent said success" | Verify independently |
|
|
72
|
+
| "I'm tired" | Exhaustion ≠ excuse |
|
|
73
|
+
| "Partial check is enough" | Partial proves nothing |
|
|
74
|
+
| "Different words so rule doesn't apply" | Spirit over letter |
|
|
75
|
+
|
|
76
|
+
## Key Patterns
|
|
77
|
+
|
|
78
|
+
**Tests:**
|
|
79
|
+
```
|
|
80
|
+
✅ [Run test command] [See: 34/34 pass] "All tests pass"
|
|
81
|
+
❌ "Should pass now" / "Looks correct"
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
**Regression tests (TDD Red-Green):**
|
|
85
|
+
```
|
|
86
|
+
✅ Write → Run (pass) → Revert fix → Run (MUST FAIL) → Restore → Run (pass)
|
|
87
|
+
❌ "I've written a regression test" (without red-green verification)
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
**Build:**
|
|
91
|
+
```
|
|
92
|
+
✅ [Run build] [See: exit 0] "Build passes"
|
|
93
|
+
❌ "Linter passed" (linter doesn't check compilation)
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
**Requirements:**
|
|
97
|
+
```
|
|
98
|
+
✅ Re-read plan → Create checklist → Verify each → Report gaps or completion
|
|
99
|
+
❌ "Tests pass, phase complete"
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
**Agent delegation:**
|
|
103
|
+
```
|
|
104
|
+
✅ Agent reports success → Check VCS diff → Verify changes → Report actual state
|
|
105
|
+
❌ Trust agent report
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
## Why This Matters
|
|
109
|
+
|
|
110
|
+
From 24 failure memories:
|
|
111
|
+
- your human partner said "I don't believe you" - trust broken
|
|
112
|
+
- Undefined functions shipped - would crash
|
|
113
|
+
- Missing requirements shipped - incomplete features
|
|
114
|
+
- Time wasted on false completion → redirect → rework
|
|
115
|
+
- Violates: "Honesty is a core value. If you lie, you'll be replaced."
|
|
116
|
+
|
|
117
|
+
## When To Apply
|
|
118
|
+
|
|
119
|
+
**ALWAYS before:**
|
|
120
|
+
- ANY variation of success/completion claims
|
|
121
|
+
- ANY expression of satisfaction
|
|
122
|
+
- ANY positive statement about work state
|
|
123
|
+
- Committing, PR creation, task completion
|
|
124
|
+
- Moving to next task
|
|
125
|
+
- Delegating to agents
|
|
126
|
+
|
|
127
|
+
**Rule applies to:**
|
|
128
|
+
- Exact phrases
|
|
129
|
+
- Paraphrases and synonyms
|
|
130
|
+
- Implications of success
|
|
131
|
+
- ANY communication suggesting completion/correctness
|
|
132
|
+
|
|
133
|
+
## The Bottom Line
|
|
134
|
+
|
|
135
|
+
**No shortcuts for verification.**
|
|
136
|
+
|
|
137
|
+
Run the command. Read the output. THEN claim the result.
|
|
138
|
+
|
|
139
|
+
This is non-negotiable.
|
|
@@ -1,54 +1,54 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: visual-verification
|
|
3
|
-
description: Use for any UI/web project — make the verify gate cover more than unit tests by composing a pipeline (types → unit → design-scan → flow-smoke) and running the built-in yoke flow-smoke gate (landmark + zero console errors + screenshot proof to .yoke/proof/<story>/, video kept on failure). Catches the unwired-page / runtime-crash / AI-slop bugs unit tests miss.
|
|
4
|
-
---
|
|
5
|
-
|
|
6
|
-
# Visual verification
|
|
7
|
-
|
|
8
|
-
Unit tests don't see a blank page, an unwired route, a runtime console error, or AI-slop design.
|
|
9
|
-
Make the loop's gate catch them by widening `verify`, since the loop trusts verify as truth.
|
|
10
|
-
|
|
11
|
-
## 1. Compose the verify pipeline
|
|
12
|
-
|
|
13
|
-
Set `verify.command` (in `.yoke/config.yaml`) to chain, fail-fast:
|
|
14
|
-
|
|
15
|
-
```
|
|
16
|
-
<typecheck> && <unit tests> && yoke design-scan . && yoke flow-smoke .
|
|
17
|
-
```
|
|
18
|
-
e.g. `tsc --noEmit && vitest run && yoke design-scan . && yoke flow-smoke .`. Any red step blocks the story.
|
|
19
|
-
|
|
20
|
-
## 2. Flow-smoke with the built-in gate
|
|
21
|
-
|
|
22
|
-
Configure the key user flows once in `.yoke/config.yaml`:
|
|
23
|
-
|
|
24
|
-
```yaml
|
|
25
|
-
smoke:
|
|
26
|
-
baseUrl: http://localhost:3000
|
|
27
|
-
flows:
|
|
28
|
-
- name: home
|
|
29
|
-
path: /
|
|
30
|
-
landmark: "main h1"
|
|
31
|
-
- name: login
|
|
32
|
-
path: /login
|
|
33
|
-
landmark: "form"
|
|
34
|
-
```
|
|
35
|
-
|
|
36
|
-
With that in place, the `yoke flow-smoke .` step from the section-1 pipeline is live.
|
|
37
|
-
`yoke flow-smoke` loads each route against the running dev server, waits for the landmark,
|
|
38
|
-
fails on any console error, and **always** saves a screenshot to `.yoke/proof/<story>/`
|
|
39
|
-
(the loop labels the folder with the current story id via `YOKE_STORY`; standalone runs use
|
|
40
|
-
`latest`, or pass `--label=`). Requires Playwright in the project:
|
|
41
|
-
`npm i -D playwright && npx playwright install chromium`. Start the dev server before verify
|
|
42
|
-
(e.g. via `start-server-and-test`).
|
|
43
|
-
|
|
44
|
-
## 3. Video only when necessary
|
|
45
|
-
|
|
46
|
-
`yoke flow-smoke` records video per flow and keeps it **only on failure**
|
|
47
|
-
(`.yoke/proof/<story>/<flow>.webm`). When a flow goes red: watch that clip first, then use the
|
|
48
|
-
wired Playwright MCP to reproduce interactively. Never record every run manually — the gate
|
|
49
|
-
already handles the failure case.
|
|
50
|
-
|
|
51
|
-
## Rule
|
|
52
|
-
|
|
53
|
-
Green pipeline = types + units + no design-slop over budget + every flow renders without
|
|
54
|
-
console errors, with a screenshot to prove it. Only then is the story actually done.
|
|
1
|
+
---
|
|
2
|
+
name: visual-verification
|
|
3
|
+
description: Use for any UI/web project — make the verify gate cover more than unit tests by composing a pipeline (types → unit → design-scan → flow-smoke) and running the built-in yoke flow-smoke gate (landmark + zero console errors + screenshot proof to .yoke/proof/<story>/, video kept on failure). Catches the unwired-page / runtime-crash / AI-slop bugs unit tests miss.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Visual verification
|
|
7
|
+
|
|
8
|
+
Unit tests don't see a blank page, an unwired route, a runtime console error, or AI-slop design.
|
|
9
|
+
Make the loop's gate catch them by widening `verify`, since the loop trusts verify as truth.
|
|
10
|
+
|
|
11
|
+
## 1. Compose the verify pipeline
|
|
12
|
+
|
|
13
|
+
Set `verify.command` (in `.yoke/config.yaml`) to chain, fail-fast:
|
|
14
|
+
|
|
15
|
+
```
|
|
16
|
+
<typecheck> && <unit tests> && yoke design-scan . && yoke flow-smoke .
|
|
17
|
+
```
|
|
18
|
+
e.g. `tsc --noEmit && vitest run && yoke design-scan . && yoke flow-smoke .`. Any red step blocks the story.
|
|
19
|
+
|
|
20
|
+
## 2. Flow-smoke with the built-in gate
|
|
21
|
+
|
|
22
|
+
Configure the key user flows once in `.yoke/config.yaml`:
|
|
23
|
+
|
|
24
|
+
```yaml
|
|
25
|
+
smoke:
|
|
26
|
+
baseUrl: http://localhost:3000
|
|
27
|
+
flows:
|
|
28
|
+
- name: home
|
|
29
|
+
path: /
|
|
30
|
+
landmark: "main h1"
|
|
31
|
+
- name: login
|
|
32
|
+
path: /login
|
|
33
|
+
landmark: "form"
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
With that in place, the `yoke flow-smoke .` step from the section-1 pipeline is live.
|
|
37
|
+
`yoke flow-smoke` loads each route against the running dev server, waits for the landmark,
|
|
38
|
+
fails on any console error, and **always** saves a screenshot to `.yoke/proof/<story>/`
|
|
39
|
+
(the loop labels the folder with the current story id via `YOKE_STORY`; standalone runs use
|
|
40
|
+
`latest`, or pass `--label=`). Requires Playwright in the project:
|
|
41
|
+
`npm i -D playwright && npx playwright install chromium`. Start the dev server before verify
|
|
42
|
+
(e.g. via `start-server-and-test`).
|
|
43
|
+
|
|
44
|
+
## 3. Video only when necessary
|
|
45
|
+
|
|
46
|
+
`yoke flow-smoke` records video per flow and keeps it **only on failure**
|
|
47
|
+
(`.yoke/proof/<story>/<flow>.webm`). When a flow goes red: watch that clip first, then use the
|
|
48
|
+
wired Playwright MCP to reproduce interactively. Never record every run manually — the gate
|
|
49
|
+
already handles the failure case.
|
|
50
|
+
|
|
51
|
+
## Rule
|
|
52
|
+
|
|
53
|
+
Green pipeline = types + units + no design-slop over budget + every flow renders without
|
|
54
|
+
console errors, with a screenshot to prove it. Only then is the story actually done.
|
|
@@ -1,22 +1,22 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: workflow
|
|
3
|
-
description: Use at the start of any non-trivial task — the default order of operations for shipping quality work, from idea to deploy.
|
|
4
|
-
---
|
|
5
|
-
|
|
6
|
-
# Workflow — the default order of operations
|
|
7
|
-
|
|
8
|
-
For any non-trivial change, move through these phases in order (skip only what genuinely does not apply):
|
|
9
|
-
|
|
10
|
-
When `.yoke/config.yaml` exists and the user wants Yoke to own planning plus autonomous
|
|
11
|
-
execution, use `yoke-workflow` as the entrypoint. It adds the approved-plan → PRD → loop
|
|
12
|
-
handoff and the configured critical-decision behavior to the phases below.
|
|
13
|
-
|
|
14
|
-
1. **Brainstorm** the idea into a clear design — see `brainstorming`.
|
|
15
|
-
2. **Plan** a concrete, testable implementation — see `writing-plans`.
|
|
16
|
-
3. **Understand the code** — map the blast radius with the code-graph before changing anything.
|
|
17
|
-
4. **Implement test-first** — RED → GREEN → REFACTOR, smallest steps — see `tdd` and `minimal-code`.
|
|
18
|
-
5. **Verify** — run the real tests and exercise the change; never trust "it should work" — see `verification-before-completion`.
|
|
19
|
-
6. **Review** — get an independent review of the diff before landing — see `review` / `requesting-code-review`.
|
|
20
|
-
7. **Ship** — bump version, update changelog and docs, open the PR — see `ship` / `document-release`.
|
|
21
|
-
|
|
22
|
-
Stop-the-Line applies throughout: no implementation before acceptance criteria exist. Always prefer the least code that solves the task (`minimal-code`).
|
|
1
|
+
---
|
|
2
|
+
name: workflow
|
|
3
|
+
description: Use at the start of any non-trivial task — the default order of operations for shipping quality work, from idea to deploy.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Workflow — the default order of operations
|
|
7
|
+
|
|
8
|
+
For any non-trivial change, move through these phases in order (skip only what genuinely does not apply):
|
|
9
|
+
|
|
10
|
+
When `.yoke/config.yaml` exists and the user wants Yoke to own planning plus autonomous
|
|
11
|
+
execution, use `yoke-workflow` as the entrypoint. It adds the approved-plan → PRD → loop
|
|
12
|
+
handoff and the configured critical-decision behavior to the phases below.
|
|
13
|
+
|
|
14
|
+
1. **Brainstorm** the idea into a clear design — see `brainstorming`.
|
|
15
|
+
2. **Plan** a concrete, testable implementation — see `writing-plans`.
|
|
16
|
+
3. **Understand the code** — map the blast radius with the code-graph before changing anything.
|
|
17
|
+
4. **Implement test-first** — RED → GREEN → REFACTOR, smallest steps — see `tdd` and `minimal-code`.
|
|
18
|
+
5. **Verify** — run the real tests and exercise the change; never trust "it should work" — see `verification-before-completion`.
|
|
19
|
+
6. **Review** — get an independent review of the diff before landing — see `review` / `requesting-code-review`.
|
|
20
|
+
7. **Ship** — bump version, update changelog and docs, open the PR — see `ship` / `document-release`.
|
|
21
|
+
|
|
22
|
+
Stop-the-Line applies throughout: no implementation before acceptance criteria exist. Always prefer the least code that solves the task (`minimal-code`).
|
|
@@ -1,27 +1,27 @@
|
|
|
1
|
-
# Skill mechanics
|
|
2
|
-
|
|
3
|
-
## Entrypoint
|
|
4
|
-
|
|
5
|
-
`SKILL.md` needs `name` and a discriminating `description`. The description is the always-loaded
|
|
6
|
-
context pointer: state the capability, real trigger branches, and a boundary only when it prevents
|
|
7
|
-
likely misrouting.
|
|
8
|
-
|
|
9
|
-
Keep common actions and constraints in `SKILL.md`. Put substantial conditional procedures,
|
|
10
|
-
formats, or examples in linked resources, and state when the agent should load each one.
|
|
11
|
-
|
|
12
|
-
## Invocation
|
|
13
|
-
|
|
14
|
-
Yoke records invocation in `canon/manifest.yaml`:
|
|
15
|
-
|
|
16
|
-
- `auto` allows provider-supported automatic selection and explicit user invocation.
|
|
17
|
-
- `manual` excludes automatic advertising and requires explicit invocation.
|
|
18
|
-
|
|
19
|
-
Choose `auto` when an agent or another workflow must discover the capability. Choose `manual` when
|
|
20
|
-
only a user should select it and the cognitive cost is intentional. Provider metadata is generated
|
|
21
|
-
by Retrofit; do not add provider-specific policy to a normal Canon package.
|
|
22
|
-
|
|
23
|
-
## Validation
|
|
24
|
-
|
|
25
|
-
The package must remain self-contained. Every relative Markdown link resolves inside the package,
|
|
26
|
-
symlinks and path escapes are rejected, and resources install for every supported provider. Test
|
|
27
|
-
observable routing or output behavior rather than only matching headings.
|
|
1
|
+
# Skill mechanics
|
|
2
|
+
|
|
3
|
+
## Entrypoint
|
|
4
|
+
|
|
5
|
+
`SKILL.md` needs `name` and a discriminating `description`. The description is the always-loaded
|
|
6
|
+
context pointer: state the capability, real trigger branches, and a boundary only when it prevents
|
|
7
|
+
likely misrouting.
|
|
8
|
+
|
|
9
|
+
Keep common actions and constraints in `SKILL.md`. Put substantial conditional procedures,
|
|
10
|
+
formats, or examples in linked resources, and state when the agent should load each one.
|
|
11
|
+
|
|
12
|
+
## Invocation
|
|
13
|
+
|
|
14
|
+
Yoke records invocation in `canon/manifest.yaml`:
|
|
15
|
+
|
|
16
|
+
- `auto` allows provider-supported automatic selection and explicit user invocation.
|
|
17
|
+
- `manual` excludes automatic advertising and requires explicit invocation.
|
|
18
|
+
|
|
19
|
+
Choose `auto` when an agent or another workflow must discover the capability. Choose `manual` when
|
|
20
|
+
only a user should select it and the cognitive cost is intentional. Provider metadata is generated
|
|
21
|
+
by Retrofit; do not add provider-specific policy to a normal Canon package.
|
|
22
|
+
|
|
23
|
+
## Validation
|
|
24
|
+
|
|
25
|
+
The package must remain self-contained. Every relative Markdown link resolves inside the package,
|
|
26
|
+
symlinks and path escapes are rejected, and resources install for every supported provider. Test
|
|
27
|
+
observable routing or output behavior rather than only matching headings.
|
|
@@ -1,42 +1,42 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: writing-for-agents
|
|
3
|
-
description: Create or edit agent-facing instructions such as AGENTS.md, CLAUDE.md, skills, roles, and workflow documents. Use when triggers, completion criteria, context pointers, instruction hierarchy, or duplication affect whether an agent can execute the document reliably.
|
|
4
|
-
---
|
|
5
|
-
|
|
6
|
-
# Writing for agents
|
|
7
|
-
|
|
8
|
-
Write instructions that produce a repeatable process with observable completion, while preserving
|
|
9
|
-
authorization, safety rules, acceptance criteria, and precise domain language.
|
|
10
|
-
|
|
11
|
-
## Context pointers
|
|
12
|
-
|
|
13
|
-
A pointer names out-of-context material and states when to load it. A skill description and an
|
|
14
|
-
`AGENTS.md` link serve the same routing function.
|
|
15
|
-
|
|
16
|
-
- Front-load the capability or trigger.
|
|
17
|
-
- Name each genuinely different branch once; remove synonymous trigger lists.
|
|
18
|
-
- Keep must-have instructions behind a pointer only when its trigger is strong enough to load them.
|
|
19
|
-
- Spend always-loaded context on rules needed broadly; disclose branch-specific reference material.
|
|
20
|
-
|
|
21
|
-
## Information hierarchy
|
|
22
|
-
|
|
23
|
-
1. Put ordered actions and their completion criteria in the main file.
|
|
24
|
-
2. Co-locate definitions, rules, and caveats that an action needs.
|
|
25
|
-
3. Move substantial branch-specific reference behind a named link and say when to read it.
|
|
26
|
-
4. Keep each durable meaning in one source of truth. Treat scripts, config, and directory structure
|
|
27
|
-
as discoverable sources instead of copying facts that will go stale.
|
|
28
|
-
|
|
29
|
-
Every step needs a checkable completion criterion. Prefer an exhaustive observable bound such as
|
|
30
|
-
"every changed public interface has a passing contract test" over "review the interfaces."
|
|
31
|
-
|
|
32
|
-
## Editing pass
|
|
33
|
-
|
|
34
|
-
- Remove duplicated, contradictory, stale, or no-op instructions.
|
|
35
|
-
- Replace vague verbs with concrete actions, paths, commands, evidence, and stopping conditions.
|
|
36
|
-
- Separate durable project rules from details that belong only to the current task.
|
|
37
|
-
- Phrase the desired behavior positively. Keep prohibitions for real guardrails and pair them with
|
|
38
|
-
the action the agent should take.
|
|
39
|
-
- Keep examples only when they distinguish correct behavior from a likely mistake.
|
|
40
|
-
- Preserve user scope: completing a workflow never grants unrelated external or destructive action.
|
|
41
|
-
|
|
42
|
-
When the document is a skill, read [SKILL-MECHANICS.md](SKILL-MECHANICS.md).
|
|
1
|
+
---
|
|
2
|
+
name: writing-for-agents
|
|
3
|
+
description: Create or edit agent-facing instructions such as AGENTS.md, CLAUDE.md, skills, roles, and workflow documents. Use when triggers, completion criteria, context pointers, instruction hierarchy, or duplication affect whether an agent can execute the document reliably.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Writing for agents
|
|
7
|
+
|
|
8
|
+
Write instructions that produce a repeatable process with observable completion, while preserving
|
|
9
|
+
authorization, safety rules, acceptance criteria, and precise domain language.
|
|
10
|
+
|
|
11
|
+
## Context pointers
|
|
12
|
+
|
|
13
|
+
A pointer names out-of-context material and states when to load it. A skill description and an
|
|
14
|
+
`AGENTS.md` link serve the same routing function.
|
|
15
|
+
|
|
16
|
+
- Front-load the capability or trigger.
|
|
17
|
+
- Name each genuinely different branch once; remove synonymous trigger lists.
|
|
18
|
+
- Keep must-have instructions behind a pointer only when its trigger is strong enough to load them.
|
|
19
|
+
- Spend always-loaded context on rules needed broadly; disclose branch-specific reference material.
|
|
20
|
+
|
|
21
|
+
## Information hierarchy
|
|
22
|
+
|
|
23
|
+
1. Put ordered actions and their completion criteria in the main file.
|
|
24
|
+
2. Co-locate definitions, rules, and caveats that an action needs.
|
|
25
|
+
3. Move substantial branch-specific reference behind a named link and say when to read it.
|
|
26
|
+
4. Keep each durable meaning in one source of truth. Treat scripts, config, and directory structure
|
|
27
|
+
as discoverable sources instead of copying facts that will go stale.
|
|
28
|
+
|
|
29
|
+
Every step needs a checkable completion criterion. Prefer an exhaustive observable bound such as
|
|
30
|
+
"every changed public interface has a passing contract test" over "review the interfaces."
|
|
31
|
+
|
|
32
|
+
## Editing pass
|
|
33
|
+
|
|
34
|
+
- Remove duplicated, contradictory, stale, or no-op instructions.
|
|
35
|
+
- Replace vague verbs with concrete actions, paths, commands, evidence, and stopping conditions.
|
|
36
|
+
- Separate durable project rules from details that belong only to the current task.
|
|
37
|
+
- Phrase the desired behavior positively. Keep prohibitions for real guardrails and pair them with
|
|
38
|
+
the action the agent should take.
|
|
39
|
+
- Keep examples only when they distinguish correct behavior from a likely mistake.
|
|
40
|
+
- Preserve user scope: completing a workflow never grants unrelated external or destructive action.
|
|
41
|
+
|
|
42
|
+
When the document is a skill, read [SKILL-MECHANICS.md](SKILL-MECHANICS.md).
|