ai-design-context 0.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +156 -0
- package/CHANGELOG.md +71 -0
- package/CODE_OF_CONDUCT.md +29 -0
- package/CONTRIBUTING.md +128 -0
- package/LICENSE +21 -0
- package/README.md +283 -0
- package/RELEASE_NOTES_v0.1.0.md +45 -0
- package/RELEASE_NOTES_v0.3.0.md +47 -0
- package/RELEASE_NOTES_v0.4.0.md +29 -0
- package/RELEASE_READINESS.md +107 -0
- package/ROADMAP.md +60 -0
- package/SECURITY.md +40 -0
- package/assets/logo.svg +9 -0
- package/benchmarks/README.md +19 -0
- package/benchmarks/chat.md +37 -0
- package/benchmarks/mobile-home-screen.md +38 -0
- package/benchmarks/notes-app.md +37 -0
- package/benchmarks/settings.md +37 -0
- package/benchmarks/shopping-list.md +37 -0
- package/benchmarks/todo-app.md +144 -0
- package/checklists/DESIGN_QA.md +102 -0
- package/checklists/GENERATED_INDEX.md +7 -0
- package/docs/AGENT_CONTEXT.md +111 -0
- package/docs/BENCHMARK.md +170 -0
- package/docs/DESIGNLINT_READINESS.md +110 -0
- package/docs/DESIGN_SYSTEM.md +23 -0
- package/docs/EVALUATION_RUBRIC.md +154 -0
- package/docs/GITHUB_SETUP.md +28 -0
- package/docs/GLOSSARY.md +45 -0
- package/docs/INDEX.md +109 -0
- package/docs/KNOWLEDGE_ENGINE.md +594 -0
- package/docs/MOBILE_FIRST.md +18 -0
- package/docs/NPM_RELEASE.md +43 -0
- package/docs/PATTERN_SPEC.md +113 -0
- package/docs/PHILOSOPHY.md +25 -0
- package/docs/PRODUCT_THINKING.md +26 -0
- package/docs/STYLE_GUIDE.md +49 -0
- package/docs/TRACEABILITY.md +45 -0
- package/evidence/README.md +64 -0
- package/evidence/TEMPLATE.md +46 -0
- package/evidence/chat/.gitkeep +1 -0
- package/evidence/mobile-home/.gitkeep +1 -0
- package/evidence/notes/.gitkeep +1 -0
- package/evidence/settings/.gitkeep +1 -0
- package/evidence/shopping/.gitkeep +1 -0
- package/evidence/todo/.gitkeep +1 -0
- package/evidence/todo/2026-06-25-codex-gpt-5/EVALUATION.md +66 -0
- package/evidence/todo/2026-06-25-codex-gpt-5/README.md +40 -0
- package/evidence/todo/2026-06-25-codex-gpt-5/ai-design-rules/generated-output.md +163 -0
- package/evidence/todo/2026-06-25-codex-gpt-5/ai-design-rules/metadata.json +38 -0
- package/evidence/todo/2026-06-25-codex-gpt-5/ai-design-rules/notes.md +20 -0
- package/evidence/todo/2026-06-25-codex-gpt-5/ai-design-rules/prompt.md +16 -0
- package/evidence/todo/2026-06-25-codex-gpt-5/ai-design-rules/scores.md +18 -0
- package/evidence/todo/2026-06-25-codex-gpt-5/baseline/generated-output.md +98 -0
- package/evidence/todo/2026-06-25-codex-gpt-5/baseline/metadata.json +19 -0
- package/evidence/todo/2026-06-25-codex-gpt-5/baseline/notes.md +19 -0
- package/evidence/todo/2026-06-25-codex-gpt-5/baseline/prompt.md +9 -0
- package/evidence/todo/2026-06-25-codex-gpt-5/baseline/scores.md +18 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/EVALUATION.md +70 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/README.md +52 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/GENERATION_NOTES.md +133 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/app.js +289 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/context/context-preserving-preview.json +443 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/context/daily-home-surface.json +373 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/context/quick-capture.json +301 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/context/sources-read.txt +36 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/generated-output.md +11 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/index.html +64 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/launch-prompt.md +1 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/metadata.json +55 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/notes.md +15 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/parent-status-message.md +1 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/prompt.md +49 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/runtime-checks.json +776 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/scores.md +20 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-completed.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-default.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-detail-reduced-motion.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-detail.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-empty.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-keyboard-focus.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-save-error.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-saved.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-saving.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-validation.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-completed.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-default.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-detail-reduced-motion.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-detail.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-empty.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-keyboard-focus.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-save-error.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-saved.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-saving.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-validation.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/source-hashes.json +5 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/styles.css +122 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/GENERATION_NOTES.md +86 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/app.js +303 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/generated-output.md +11 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/index.html +71 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/launch-prompt.md +1 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/metadata.json +55 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/notes.md +15 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/parent-status-message.md +1 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/prompt.md +49 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/runtime-checks.json +772 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/scores.md +20 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-completed.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-default.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-detail-reduced-motion.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-detail.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-empty.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-keyboard-focus.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-save-error.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-saved.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-saving.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-validation.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-completed.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-default.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-detail-reduced-motion.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-detail.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-empty.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-keyboard-focus.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-save-error.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-saved.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-saving.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-validation.png +0 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/source-hashes.json +5 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/baseline/styles.css +9 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/capture-ai-design-rules.js +60 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/capture-baseline.js +60 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/common/brief.md +35 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/common/knowledge-files.json +127 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/common/knowledge.patch +890 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/common/review-context.json +550 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/common/seed.json +86 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/common/setup.json +47 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/common/technical-envelope.md +11 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/layout-metrics.json +133 -0
- package/evidence/todo/2026-09-28-codex-paired-todo/layout-probe.js +15 -0
- package/examples/GENERATED_INDEX.md +8 -0
- package/examples/INDEX.md +16 -0
- package/examples/README.md +15 -0
- package/examples/consequential-action-confirmation-reference.md +83 -0
- package/examples/todo-app-benchmark.md +117 -0
- package/examples/todo-reference/README.md +29 -0
- package/examples/todo-reference/app.js +154 -0
- package/examples/todo-reference/favicon.svg +4 -0
- package/examples/todo-reference/index.html +67 -0
- package/examples/todo-reference/review-evidence/2026-07-15/desktop-default.png +0 -0
- package/examples/todo-reference/review-evidence/2026-07-15/desktop-detail.png +0 -0
- package/examples/todo-reference/review-evidence/2026-07-15/mobile-detail.png +0 -0
- package/examples/todo-reference/review-evidence/2026-07-15/mobile-error.png +0 -0
- package/examples/todo-reference/styles.css +241 -0
- package/graph/GENERATED_GRAPH.md +71 -0
- package/observations/README.md +23 -0
- package/observations/accessibility/reduced-interaction-motion.md +49 -0
- package/observations/accessibility/textual-error-recovery.md +50 -0
- package/observations/accessibility/visible-keyboard-focus.md +49 -0
- package/observations/ai/agent-consequential-action-controls.md +56 -0
- package/observations/ai/agent-work-visibility-and-intervention.md +51 -0
- package/observations/ai/figma-agent-context-and-skills.md +52 -0
- package/observations/finance/high-risk-wallet-recovery-flows.md +55 -0
- package/observations/finance/market-interface-adaptive-trading-surfaces.md +55 -0
- package/observations/interaction/chrome-scoped-view-transitions.md +52 -0
- package/observations/interaction/figma-motion-as-system-capability.md +53 -0
- package/observations/mobile/mobile-bottom-sheet-choice-contracts.md +55 -0
- package/observations/performance/reserved-loading-space.md +51 -0
- package/observations/product-repos/pinned-ai-design-rules-integration.md +49 -0
- package/observations/visual/apple-liquid-glass-functional-layer.md +52 -0
- package/observations/visual/chrome-content-first-adaptive-ui.md +51 -0
- package/observations/visual/material-3-expressive-system.md +52 -0
- package/package.json +65 -0
- package/patterns/GENERATED_INDEX.md +13 -0
- package/patterns/INDEX.md +41 -0
- package/patterns/README.md +5 -0
- package/patterns/consequential-action-confirmation.md +127 -0
- package/patterns/context-preserving-preview.md +145 -0
- package/patterns/daily-home-surface.md +134 -0
- package/patterns/mobile-primary-action.md +110 -0
- package/patterns/object-status-list.md +114 -0
- package/patterns/progressive-detail.md +113 -0
- package/patterns/quick-capture.md +132 -0
- package/prompts/CONSEQUENTIAL_ACTION_REVIEW.md +100 -0
- package/prompts/DESIGN_EVIDENCE_REVIEW.md +49 -0
- package/prompts/GENERATED_INDEX.md +11 -0
- package/prompts/INDEX.md +24 -0
- package/prompts/PROTOTYPE_REVIEW.md +189 -0
- package/prompts/QUICK_CAPTURE_STATE_REVIEW.md +74 -0
- package/prompts/TODO_APP_BENCHMARK.md +88 -0
- package/registry/objects.json +568 -0
- package/registry/relationships.json +972 -0
- package/research/GENERATED_INDEX.md +18 -0
- package/research/INDEX.md +40 -0
- package/research/accessibility/keyboard-focus-and-interaction-motion.md +62 -0
- package/research/accessibility/textual-error-recovery.md +57 -0
- package/research/performance/reserved-loading-space.md +56 -0
- package/research/products/apple-reminders.md +59 -0
- package/research/products/arc.md +61 -0
- package/research/products/linear.md +62 -0
- package/research/products/telegram.md +61 -0
- package/research/products/things-3.md +60 -0
- package/research/ux/contextual-agentic-interface.md +68 -0
- package/research/ux/high-risk-operational-flows.md +89 -0
- package/research/ux/motion-as-state-continuity.md +71 -0
- package/research/visual/expressive-system-ui-2026.md +75 -0
- package/reviews/GENERATED_INDEX.md +9 -0
- package/reviews/README.md +25 -0
- package/reviews/consequential-action-confirmation-review.md +92 -0
- package/reviews/todo-app-benchmark-reference-review.md +86 -0
- package/reviews/todo-reference-fixture-review.md +77 -0
- package/rules/GENERATED_INDEX.md +23 -0
- package/rules/INDEX.md +37 -0
- package/rules/accessibility/A11Y-001.md +58 -0
- package/rules/accessibility/A11Y-002.md +55 -0
- package/rules/accessibility/A11Y-003.md +54 -0
- package/rules/accessibility/A11Y-004.md +54 -0
- package/rules/ia/IA-001.md +60 -0
- package/rules/ia/IA-002.md +61 -0
- package/rules/performance/PERF-001.md +55 -0
- package/rules/product/PRD-001.md +61 -0
- package/rules/product/PRD-002.md +61 -0
- package/rules/ux/UX-001.md +58 -0
- package/rules/ux/UX-002.md +61 -0
- package/rules/ux/UX-003.md +60 -0
- package/rules/ux/UX-004.md +61 -0
- package/rules/ux/UX-005.md +62 -0
- package/rules/ux/UX-006.md +69 -0
- package/rules/visual/VIS-001.md +59 -0
- package/rules/visual/VIS-002.md +59 -0
- package/schema/checklist.schema.json +21 -0
- package/schema/common.schema.json +178 -0
- package/schema/observation.schema.json +21 -0
- package/schema/pattern.schema.json +21 -0
- package/schema/prompt.schema.json +21 -0
- package/schema/reference-project.schema.json +21 -0
- package/schema/research.schema.json +21 -0
- package/schema/review.schema.json +21 -0
- package/schema/rule.schema.json +21 -0
- package/skills/README.md +48 -0
- package/skills/accessibility-reviewer/SKILL.md +101 -0
- package/skills/agent-context/SKILL.md +34 -0
- package/skills/agent-context/agents/openai.yaml +4 -0
- package/skills/design-evidence-researcher/SKILL.md +47 -0
- package/skills/design-evidence-researcher/agents/openai.yaml +4 -0
- package/skills/design-reviewer/SKILL.md +113 -0
- package/skills/design-system-architect/SKILL.md +101 -0
- package/skills/information-architect/SKILL.md +97 -0
- package/skills/interaction-designer/SKILL.md +104 -0
- package/skills/knowledge-graph-architect/SKILL.md +96 -0
- package/skills/mobile-ux-expert/SKILL.md +105 -0
- package/skills/motion-designer/SKILL.md +100 -0
- package/skills/performance-reviewer/SKILL.md +97 -0
- package/skills/product-designer/SKILL.md +104 -0
- package/skills/prompt-architect/SKILL.md +105 -0
- package/skills/reference-driven-design/SKILL.md +119 -0
- package/skills/ux-reviewer/SKILL.md +99 -0
- package/skills/visual-designer/SKILL.md +107 -0
- package/starter-kit/AGENTS.md +37 -0
- package/starter-kit/BOOTSTRAP.md +13 -0
- package/starter-kit/INSTALL_WITH_AGENT.md +133 -0
- package/starter-kit/PROJECT_INTEGRATION.md +60 -0
- package/starter-kit/README.md +41 -0
- package/starter-kit/benchmarks/BENCHMARK_CHECKLIST.md +10 -0
- package/starter-kit/docs/DESIGN_DECISIONS.md +22 -0
- package/starter-kit/docs/INFORMATION_ARCHITECTURE.md +25 -0
- package/starter-kit/docs/PERSONAS.md +14 -0
- package/starter-kit/docs/PRD.md +32 -0
- package/starter-kit/docs/USER_FLOWS.md +19 -0
- package/starter-kit/reviews/DESIGN_REVIEW.md +49 -0
- package/starter-kit/templates/FEATURE_TEMPLATE.md +36 -0
- package/starter-kit/templates/TASK_TEMPLATE.md +22 -0
- package/templates/CHECKLIST_TEMPLATE.md +43 -0
- package/templates/OBSERVATION_TEMPLATE.md +41 -0
- package/templates/PATTERN_TEMPLATE.md +100 -0
- package/templates/PROMPT_TEMPLATE.md +46 -0
- package/templates/PROTOTYPE_REVIEW.md +50 -0
- package/templates/REFERENCE_PROJECT_TEMPLATE.md +49 -0
- package/templates/REVIEW_TEMPLATE.md +55 -0
- package/templates/RULE_TEMPLATE.md +57 -0
- package/tools/cli.mjs +46 -0
- package/tools/context.mjs +342 -0
- package/tools/designlint.mjs +77 -0
- package/tools/generate-indexes.mjs +621 -0
- package/tools/init.mjs +99 -0
- package/tools/validate-benchmark-evidence.mjs +248 -0
- package/tools/validate-knowledge.mjs +481 -0
|
@@ -0,0 +1,144 @@
|
|
|
1
|
+
# Todo App Benchmark
|
|
2
|
+
|
|
3
|
+
This is the reference benchmark scenario.
|
|
4
|
+
|
|
5
|
+
Do not generate comparison results in this file.
|
|
6
|
+
|
|
7
|
+
## Product Brief
|
|
8
|
+
|
|
9
|
+
Create a modern consumer todo app for everyday personal task management.
|
|
10
|
+
|
|
11
|
+
The app should help a busy person capture tasks quickly, see what matters today, complete tasks, and inspect optional details without turning the product into a project management dashboard.
|
|
12
|
+
|
|
13
|
+
## Expected User
|
|
14
|
+
|
|
15
|
+
- Adult consumer managing personal errands, reminders, and daily tasks.
|
|
16
|
+
- Uses mobile frequently and desktop occasionally.
|
|
17
|
+
- Wants low friction more than advanced planning.
|
|
18
|
+
|
|
19
|
+
## Expected Platform
|
|
20
|
+
|
|
21
|
+
- Responsive web app.
|
|
22
|
+
- Must work well at mobile width around 390px.
|
|
23
|
+
- Desktop layout may use additional space but must preserve the same product model.
|
|
24
|
+
|
|
25
|
+
## Constraints
|
|
26
|
+
|
|
27
|
+
- No authentication flow required.
|
|
28
|
+
- No team collaboration required.
|
|
29
|
+
- No analytics dashboard.
|
|
30
|
+
- No project management concepts such as sprints, roadmaps, or velocity.
|
|
31
|
+
- Include realistic empty, loading, and error states.
|
|
32
|
+
- Include accessible labels and touch-friendly controls.
|
|
33
|
+
- Keep optional metadata secondary.
|
|
34
|
+
|
|
35
|
+
## Core User Tasks
|
|
36
|
+
|
|
37
|
+
- Add a task quickly.
|
|
38
|
+
- See today's tasks.
|
|
39
|
+
- Mark a task complete.
|
|
40
|
+
- Inspect or edit task details.
|
|
41
|
+
- Recover from a failed save.
|
|
42
|
+
|
|
43
|
+
## Evaluation Criteria
|
|
44
|
+
|
|
45
|
+
Use `docs/EVALUATION_RUBRIC.md`.
|
|
46
|
+
|
|
47
|
+
Pay special attention to:
|
|
48
|
+
|
|
49
|
+
- Product Thinking;
|
|
50
|
+
- UX;
|
|
51
|
+
- Information Architecture;
|
|
52
|
+
- Mobile-first;
|
|
53
|
+
- State Design;
|
|
54
|
+
- Simplicity;
|
|
55
|
+
- Overall Product Quality.
|
|
56
|
+
|
|
57
|
+
## Baseline Prompt
|
|
58
|
+
|
|
59
|
+
```text
|
|
60
|
+
Design and implement a responsive todo app for everyday personal task management.
|
|
61
|
+
|
|
62
|
+
The app should let a user add tasks, see today's tasks, complete tasks, and edit details.
|
|
63
|
+
|
|
64
|
+
Make it modern, usable, and polished.
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
## AI Design Context Prompt
|
|
68
|
+
|
|
69
|
+
```text
|
|
70
|
+
Design and implement a responsive todo app for everyday personal task management.
|
|
71
|
+
|
|
72
|
+
Use AI Design Context from this repository.
|
|
73
|
+
|
|
74
|
+
Follow the knowledge graph:
|
|
75
|
+
research -> rules -> patterns -> prompts -> reference projects -> reviews
|
|
76
|
+
|
|
77
|
+
Use existing rules and patterns only. Do not invent new rules or patterns.
|
|
78
|
+
|
|
79
|
+
The app should let a user add tasks, see today's tasks, complete tasks, and edit details.
|
|
80
|
+
|
|
81
|
+
Prioritize low-friction capture, a stable daily home surface, mobile-first behavior, context-preserving detail, accessible touch targets, and clear empty/loading/error states.
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
## Expected Outputs
|
|
85
|
+
|
|
86
|
+
Both benchmark runs should publish:
|
|
87
|
+
|
|
88
|
+
- prompt used;
|
|
89
|
+
- generated source code or artifact;
|
|
90
|
+
- screenshots or preview links for mobile and desktop;
|
|
91
|
+
- notes about tool access and model settings;
|
|
92
|
+
- evaluation table using `docs/EVALUATION_RUBRIC.md`.
|
|
93
|
+
|
|
94
|
+
The expected product output should include:
|
|
95
|
+
|
|
96
|
+
- daily task home surface;
|
|
97
|
+
- quick add task path;
|
|
98
|
+
- completed task behavior;
|
|
99
|
+
- detail inspection or edit behavior;
|
|
100
|
+
- empty state;
|
|
101
|
+
- loading state;
|
|
102
|
+
- error state;
|
|
103
|
+
- accessible labels;
|
|
104
|
+
- mobile layout.
|
|
105
|
+
|
|
106
|
+
## Evaluation Template
|
|
107
|
+
|
|
108
|
+
```md
|
|
109
|
+
# Todo App Evaluation
|
|
110
|
+
|
|
111
|
+
Scenario: Todo App
|
|
112
|
+
Model:
|
|
113
|
+
Date:
|
|
114
|
+
Evaluator:
|
|
115
|
+
|
|
116
|
+
| Category | Baseline | AI Design Context | Notes |
|
|
117
|
+
| --- | ---: | ---: | --- |
|
|
118
|
+
| Product Thinking | | | |
|
|
119
|
+
| UX | | | |
|
|
120
|
+
| Information Architecture | | | |
|
|
121
|
+
| Navigation | | | |
|
|
122
|
+
| Accessibility | | | |
|
|
123
|
+
| Mobile-first | | | |
|
|
124
|
+
| Visual Hierarchy | | | |
|
|
125
|
+
| State Design | | | |
|
|
126
|
+
| Consistency | | | |
|
|
127
|
+
| Performance Awareness | | | |
|
|
128
|
+
| Simplicity | | | |
|
|
129
|
+
| Overall Product Quality | | | |
|
|
130
|
+
|
|
131
|
+
Baseline average:
|
|
132
|
+
AI Design Context average:
|
|
133
|
+
Difference:
|
|
134
|
+
|
|
135
|
+
## Evidence Links
|
|
136
|
+
|
|
137
|
+
- Baseline output:
|
|
138
|
+
- AI Design Context output:
|
|
139
|
+
- Mobile screenshots:
|
|
140
|
+
- Desktop screenshots:
|
|
141
|
+
|
|
142
|
+
## Notes
|
|
143
|
+
|
|
144
|
+
```
|
|
@@ -0,0 +1,102 @@
|
|
|
1
|
+
---
|
|
2
|
+
id: CHECK-00001
|
|
3
|
+
alias: CHECK-DESIGN-QA
|
|
4
|
+
slug: design-qa
|
|
5
|
+
title: Design QA Checklist
|
|
6
|
+
object_type: checklist
|
|
7
|
+
status: draft
|
|
8
|
+
version: 0.4.0
|
|
9
|
+
category: design-qa
|
|
10
|
+
tags:
|
|
11
|
+
- checklist
|
|
12
|
+
- review
|
|
13
|
+
- design-qa
|
|
14
|
+
maturity: seed
|
|
15
|
+
risk_level: medium
|
|
16
|
+
relationships:
|
|
17
|
+
- type: related_to
|
|
18
|
+
target: RULE-00001
|
|
19
|
+
- type: related_to
|
|
20
|
+
target: RULE-00002
|
|
21
|
+
- type: related_to
|
|
22
|
+
target: RULE-00003
|
|
23
|
+
- type: related_to
|
|
24
|
+
target: RULE-00004
|
|
25
|
+
- type: related_to
|
|
26
|
+
target: RULE-00005
|
|
27
|
+
- type: related_to
|
|
28
|
+
target: RULE-00006
|
|
29
|
+
- type: related_to
|
|
30
|
+
target: RULE-00007
|
|
31
|
+
- type: related_to
|
|
32
|
+
target: RULE-00008
|
|
33
|
+
- type: related_to
|
|
34
|
+
target: RULE-00009
|
|
35
|
+
- type: related_to
|
|
36
|
+
target: RULE-00010
|
|
37
|
+
- type: related_to
|
|
38
|
+
target: RULE-00011
|
|
39
|
+
- type: related_to
|
|
40
|
+
target: RULE-00012
|
|
41
|
+
- type: related_to
|
|
42
|
+
target: RULE-00013
|
|
43
|
+
- type: related_to
|
|
44
|
+
target: RULE-00014
|
|
45
|
+
- type: related_to
|
|
46
|
+
target: RULE-00015
|
|
47
|
+
- type: related_to
|
|
48
|
+
target: RULE-00016
|
|
49
|
+
- type: related_to
|
|
50
|
+
target: RULE-00017
|
|
51
|
+
---
|
|
52
|
+
|
|
53
|
+
# Design QA Checklist
|
|
54
|
+
|
|
55
|
+
Use this checklist after applying research, rules, and patterns.
|
|
56
|
+
|
|
57
|
+
## Rule Coverage
|
|
58
|
+
|
|
59
|
+
- `PRD-001` - Is the repeated daily action prioritized?
|
|
60
|
+
- `PRD-002` - Can users capture before organizing?
|
|
61
|
+
- `IA-001` - Is there one stable home surface?
|
|
62
|
+
- `IA-002` - Are core objects, states, and relationships clear?
|
|
63
|
+
- `UX-001` - Are primary mobile actions thumb-reachable?
|
|
64
|
+
- `UX-002` - Are advanced controls progressively disclosed?
|
|
65
|
+
- `UX-003` - Can users inspect without losing context?
|
|
66
|
+
- `VIS-001` - Are semantic tokens used for visual decisions?
|
|
67
|
+
- `A11Y-001` - Are touch targets at least 44x44 CSS pixels?
|
|
68
|
+
- `A11Y-002` - Are validation errors specific, textual, and recoverable?
|
|
69
|
+
- `A11Y-003` - Does every keyboard-operable control retain visible focus?
|
|
70
|
+
- `A11Y-004` - When non-essential motion exists, can users reduce or disable it?
|
|
71
|
+
- `PERF-001` - Does asynchronous content preserve surrounding layout geometry?
|
|
72
|
+
- `VIS-002` - Do expressive materials preserve content and primary-action priority?
|
|
73
|
+
- `UX-004` - Does motion explain state continuity with reduced-motion and platform fallbacks?
|
|
74
|
+
- `UX-005` - When work is delegated to an agent, are status, result, and intervention points observable?
|
|
75
|
+
- `UX-006` - Are consequential actions confirmed with preserved object, consequence, and recovery context?
|
|
76
|
+
|
|
77
|
+
## Pattern Coverage
|
|
78
|
+
|
|
79
|
+
- Does the screen use a documented pattern when one applies?
|
|
80
|
+
- Is the pattern appropriate for the product context?
|
|
81
|
+
- Did the pattern preserve the user's primary job?
|
|
82
|
+
- Did the pattern avoid copying a source product blindly?
|
|
83
|
+
- If consequential side effects exist, did the screen use or explicitly reject `PAT-007`?
|
|
84
|
+
|
|
85
|
+
## Final Checks
|
|
86
|
+
|
|
87
|
+
- Product goal is clear.
|
|
88
|
+
- Primary and secondary actions are visually distinct.
|
|
89
|
+
- Empty, loading, error, success, and disabled states are accounted for.
|
|
90
|
+
- Mobile layout works at 390px width.
|
|
91
|
+
- Accessibility basics are present: labels, focus, contrast, target size, and textual error recovery.
|
|
92
|
+
- State changes avoid unexpected layout shifts.
|
|
93
|
+
- Motion, when present, has a reduced-motion equivalent.
|
|
94
|
+
- Platform-specific visual effects and interaction APIs have appropriate fallbacks.
|
|
95
|
+
- Visual direction cites observable references instead of imitating a brand.
|
|
96
|
+
- Consequential actions preserve critical context through confirmation, challenge, error, retry, and success states.
|
|
97
|
+
- Before broad design work, the agent asked whether to show multiple options or proceed with one conservative implementation path.
|
|
98
|
+
|
|
99
|
+
## Evidence Boundary
|
|
100
|
+
|
|
101
|
+
- Mark unrendered, unmeasured, or untested behavior as a gap rather than a pass.
|
|
102
|
+
- Cite the reviewed artifact, prompt, reference project, or benchmark evidence in the resulting review object.
|
|
@@ -0,0 +1,7 @@
|
|
|
1
|
+
# Generated Checklists Index
|
|
2
|
+
|
|
3
|
+
This file is generated from registry metadata. Do not edit manually.
|
|
4
|
+
|
|
5
|
+
| ID | Alias | Title | Category | Status | Maturity | File Path | Rule Coverage | Used By |
|
|
6
|
+
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
|
|
7
|
+
| CHECK-00001 | CHECK-DESIGN-QA | Design QA Checklist | design-qa | draft | seed | checklists/DESIGN_QA.md | related_to: PRD-001 / RULE-00001 Daily Actions First<br>related_to: PRD-002 / RULE-00002 Low-Friction Capture Before Organization<br>related_to: IA-001 / RULE-00003 One Stable Home Surface<br>related_to: IA-002 / RULE-00004 Define The Object Model Before Screens<br>related_to: UX-001 / RULE-00005 Thumb-Zone Primary Actions<br>related_to: UX-002 / RULE-00006 Progressively Disclose Power<br>related_to: UX-003 / RULE-00007 Preserve Context During Inspection<br>related_to: VIS-001 / RULE-00008 Semantic Tokens Only<br>related_to: A11Y-001 / RULE-00009 44x44 Touch Targets<br>related_to: A11Y-002 / RULE-00010 Textual Input Error Recovery<br>related_to: PERF-001 / RULE-00011 Reserve Layout Space For Asynchronous Content<br>related_to: A11Y-003 / RULE-00012 Visible Keyboard Focus<br>related_to: A11Y-004 / RULE-00013 Reduce Non-Essential Interaction Motion<br>related_to: VIS-002 / RULE-00014 Keep Expression Subordinate To Content<br>related_to: UX-004 / RULE-00015 Use Motion To Explain State Continuity<br>related_to: UX-005 / RULE-00016 Expose Agent Work And Intervention<br>related_to: UX-006 / RULE-00017 Confirm Consequential Actions With Preserved Context | requires: PROMPT-DESIGN-EVIDENCE-REVIEW / PROMPT-00005 Design Evidence Review Prompt<br>requires: REVIEW-TODO-BENCHMARK / REVIEW-00001 Todo App Benchmark Reference Review<br>requires: REVIEW-TODO-RUNNABLE-FIXTURE / REVIEW-00002 Todo Runnable Fixture Review<br>requires: PROMPT-PROTOTYPE-REVIEW / PROMPT-00002 Prototype Review Prompt<br>requires: REVIEW-CONSEQUENTIAL-ACTION-CONFIRMATION / REVIEW-00003 Consequential Action Confirmation Validation Review |
|
|
@@ -0,0 +1,111 @@
|
|
|
1
|
+
# Agent Context CLI
|
|
2
|
+
|
|
3
|
+
`npm run context` converts a task, review target, or known graph object into a compact, traceable context bundle. It is a read-only local tool: no network, authentication, or registry writes are involved.
|
|
4
|
+
|
|
5
|
+
## Installed npm Package
|
|
6
|
+
|
|
7
|
+
After installing `ai-design-context` as a dev dependency, run the packaged CLI from the product repository:
|
|
8
|
+
|
|
9
|
+
```bash
|
|
10
|
+
npx --no-install ai-design-context context --task quick-capture --platform mobile --intent implement
|
|
11
|
+
npx --no-install ai-design-context context --review REF-00001 --intent qa
|
|
12
|
+
npx --no-install ai-design-context context --object PAT-00002 --format json
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
The installed CLI always reads its own graph, not a `registry/` directory in the consuming project. Markdown lists absolute paths so the agent can open the selected files. JSON preserves the relative graph `path`, adds an `absolutePath` to anchors and objects, and adds the package's `knowledgeRoot` at the top level. These reading paths depend on the installation location and must not be used as stable object identifiers.
|
|
16
|
+
|
|
17
|
+
`npx ai-design-context init` appends a marked block to `AGENTS.md` and creates only missing starter-kit files. It preserves current project instructions, populated files, and existing marked blocks. It rejects malformed markers, symlinked destinations, and incompatible file/directory destinations before creating templates. Installation itself performs no initialization.
|
|
18
|
+
|
|
19
|
+
Retrieval is lexical and read-only. A `--review` query matches a known graph object or phrase; it does not inspect or judge arbitrary files in the product repository. Resolve relevant context, then inspect the implementation separately. See [installation](../starter-kit/INSTALL_WITH_AGENT.md) for pre-publication tarball use.
|
|
20
|
+
|
|
21
|
+
## Commands
|
|
22
|
+
|
|
23
|
+
```bash
|
|
24
|
+
# Design or implement a feature.
|
|
25
|
+
npm run context -- --task quick-capture --platform mobile --intent implement
|
|
26
|
+
npm run context -- --task "build a shopping list" --intent implement
|
|
27
|
+
|
|
28
|
+
# Review an implementation, benchmark artifact, or reference fixture.
|
|
29
|
+
npm run context -- --review examples/todo-reference --intent qa
|
|
30
|
+
|
|
31
|
+
# Resolve an exact stable object.
|
|
32
|
+
npm run context -- --object PAT-00002 --format json
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
Exactly one of `--task`, `--review`, or `--object` is required.
|
|
36
|
+
|
|
37
|
+
## Output Contract
|
|
38
|
+
|
|
39
|
+
Markdown is the human default. `--format json` emits only this stable shape on stdout:
|
|
40
|
+
|
|
41
|
+
```json
|
|
42
|
+
{
|
|
43
|
+
"query": { "mode": "task", "value": "quick-capture", "platform": "mobile", "intent": "implement" },
|
|
44
|
+
"anchors": [],
|
|
45
|
+
"objects": []
|
|
46
|
+
}
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
Every entry in `anchors` and `objects` retains `id`, `alias`, `title`, `type`, `status`, `maturity`, and `path`, and now adds a `reason` object. Consumers should allow additive fields. For example:
|
|
50
|
+
|
|
51
|
+
```json
|
|
52
|
+
{
|
|
53
|
+
"id": "RULE-00002",
|
|
54
|
+
"alias": "PRD-002",
|
|
55
|
+
"title": "Low-Friction Capture Before Organization",
|
|
56
|
+
"type": "rule",
|
|
57
|
+
"status": "draft",
|
|
58
|
+
"maturity": "seed",
|
|
59
|
+
"path": "rules/product/PRD-002.md",
|
|
60
|
+
"reason": { "kind": "dependency", "source": "PAT-00002", "relationship": "requires" }
|
|
61
|
+
}
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
## Matching And Scope
|
|
65
|
+
|
|
66
|
+
Task and review modes search stable registry fields and object Markdown. An exact ID, alias, slug, or registry path wins first. Otherwise, whole Unicode words must match; titles rank above slugs, aliases, categories, and incidental body mentions. Shorter focused titles break otherwise equal title matches. One highest-ranked anchor is selected, with stable ID ordering as the final tie-breaker.
|
|
67
|
+
|
|
68
|
+
The small English filler set (`a`, `an`, `the`, `please`, `build`, `create`, `implement`, `make`, `for`, `me`) is ignored in non-exact queries. Thus `build a shopping list` and `shopping list` resolve alike. Unknown subject terms are not dropped. This is lexical retrieval, not translation or semantic search: use a known slug or ID when a paraphrase or another language has no matching words. A query containing only fillers fails instead of selecting arbitrary knowledge.
|
|
69
|
+
|
|
70
|
+
Object mode matches only an exact ID, alias, slug, or registry path and does not apply filler handling.
|
|
71
|
+
|
|
72
|
+
From the anchor the resolver includes:
|
|
73
|
+
|
|
74
|
+
- Direct `related_to` neighbors for research, rules, and patterns, without recursively following optional neighbors.
|
|
75
|
+
- Rules derived from a research anchor through `derived_from` or `inspired_by`.
|
|
76
|
+
- Direct review prompts and checklists for anchors and those derived rules; reviews that `validate` a reference-project anchor.
|
|
77
|
+
- All transitive registered `requires`, `derived_from`, `inspired_by`, and `implements` dependencies of every selected object; `validates` targets of selected reviews and reference projects.
|
|
78
|
+
|
|
79
|
+
Dependencies are followed to completion even across cycles. A missing registered dependency is an error. The resolver does not walk backward from every shared rule into every consuming pattern or generation prompt. Broad review checklists remain available as gates, but their optional `related_to` links do not automatically select the whole graph. Follow those links explicitly when broadening a review. Non-registry observation sources remain in the research files and must be read there.
|
|
80
|
+
|
|
81
|
+
Platform matching is additive. It selects at most one matching rule or pattern from eligible types **before** ranking, then includes its dependencies and direct context under the same rules. A platform with no match adds nothing. Platform matching does not replace the task anchor or prove applicability.
|
|
82
|
+
|
|
83
|
+
There is no hard object-count cap: completeness of required dependencies takes priority over a fixed limit. Large mandatory graphs may still produce large bundles.
|
|
84
|
+
|
|
85
|
+
## Selection Reasons
|
|
86
|
+
|
|
87
|
+
Markdown explains the selection beside each object. JSON records the first deterministic inclusion reason:
|
|
88
|
+
|
|
89
|
+
| `reason.kind` | Meaning | Additional fields |
|
|
90
|
+
| --- | --- | --- |
|
|
91
|
+
| `match` | Exact match or highest-ranked lexical match | `field`, `terms` matched in that field |
|
|
92
|
+
| `platform` | Added by the platform query | `value`, `field`, `terms` |
|
|
93
|
+
| `dependency` | Outgoing mandatory or validation edge from a selected object | `source`, `relationship` |
|
|
94
|
+
| `related` | Direct optional neighbor of a seed object | `source`, `relationship` |
|
|
95
|
+
| `derived_rule` | Rule points back to selected research | `source` (research), `relationship` |
|
|
96
|
+
| `review` | Review gate points back to the selected object | `source` (reviewed object), `relationship` |
|
|
97
|
+
|
|
98
|
+
For `derived_rule` and `review`, the stored graph edge points from the returned object to `reason.source`. For `dependency` and `related`, it points from `reason.source` to the returned object. A reason explains inclusion; it is not a confidence score or a claim of validation.
|
|
99
|
+
|
|
100
|
+
`--intent implement` gives an implementation-first use order. `--intent qa` gives an evidence-and-checklist-first use order. Task mode defaults to `implement`; review mode defaults to `qa`.
|
|
101
|
+
|
|
102
|
+
An unmatched query returns a non-zero exit code; with `--format json` the error is a JSON object on stderr.
|
|
103
|
+
|
|
104
|
+
## Agent Use
|
|
105
|
+
|
|
106
|
+
1. Resolve context before proposing product or UI changes.
|
|
107
|
+
2. Read the listed research and rules before applying a pattern or prompt.
|
|
108
|
+
3. Treat `draft` and `seed` objects as guidance with explicit limits, not as validated truth.
|
|
109
|
+
4. Use listed checklists and reviews as quality gates.
|
|
110
|
+
5. If no context matches, capture an observation or research need; do not invent a rule to fill the gap.
|
|
111
|
+
6. Use `reason` to distinguish task matches, required knowledge, and review gates. Apply only guidance relevant to the current task.
|
|
@@ -0,0 +1,170 @@
|
|
|
1
|
+
# Benchmark
|
|
2
|
+
|
|
3
|
+
AI Design Context must be evaluated by generated product quality, not by how complete the documentation looks.
|
|
4
|
+
|
|
5
|
+
This benchmark compares two outputs for the same product brief:
|
|
6
|
+
|
|
7
|
+
```text
|
|
8
|
+
Baseline AI -> AI + AI Design Context
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
## Purpose
|
|
12
|
+
|
|
13
|
+
The benchmark exists to answer one question:
|
|
14
|
+
|
|
15
|
+
Does using AI Design Context produce better AI-generated consumer product experiences than using the same AI system without this repository?
|
|
16
|
+
|
|
17
|
+
The benchmark does not prove universal design quality. It tests whether the repository improves output on repeatable product scenarios.
|
|
18
|
+
|
|
19
|
+
## Evaluation Philosophy
|
|
20
|
+
|
|
21
|
+
Evaluation should be:
|
|
22
|
+
|
|
23
|
+
- comparative: score baseline and AI Design Context outputs side by side;
|
|
24
|
+
- reproducible: use the same model, brief, temperature, and output target;
|
|
25
|
+
- evidence-based: publish prompts, outputs, screenshots or code, and scores;
|
|
26
|
+
- conservative: do not claim improvement without measured results;
|
|
27
|
+
- practical: score product quality, not adherence to repository wording.
|
|
28
|
+
|
|
29
|
+
The reviewer should not reward an output for mentioning AI Design Context. Reward only visible product quality.
|
|
30
|
+
|
|
31
|
+
## Benchmark Process
|
|
32
|
+
|
|
33
|
+
1. Choose one benchmark scenario from `benchmarks/`.
|
|
34
|
+
2. Select one model and one generation surface.
|
|
35
|
+
3. Generate the baseline output using only the scenario brief.
|
|
36
|
+
4. Generate the AI Design Context output using the same scenario brief plus the repository instructions.
|
|
37
|
+
5. Keep model, temperature, context window, tool access, and implementation target the same.
|
|
38
|
+
6. Review both outputs using `docs/EVALUATION_RUBRIC.md`.
|
|
39
|
+
7. Store prompts, outputs, screenshots or links, evaluator notes, and scores in `evidence/`.
|
|
40
|
+
8. Declare `evidence_level` as `directional` or `rendered` in both run metadata files, then run `npm run benchmark:validate`.
|
|
41
|
+
9. Report aggregate score and per-category differences.
|
|
42
|
+
|
|
43
|
+
## Required Run Metadata
|
|
44
|
+
|
|
45
|
+
Each benchmark run must record:
|
|
46
|
+
|
|
47
|
+
- benchmark scenario;
|
|
48
|
+
- date;
|
|
49
|
+
- model name;
|
|
50
|
+
- model provider;
|
|
51
|
+
- generation surface;
|
|
52
|
+
- temperature or sampling settings;
|
|
53
|
+
- tools available;
|
|
54
|
+
- implementation target;
|
|
55
|
+
- baseline prompt;
|
|
56
|
+
- AI Design Context prompt;
|
|
57
|
+
- evaluator name or handle;
|
|
58
|
+
- rubric version;
|
|
59
|
+
- links to generated outputs.
|
|
60
|
+
|
|
61
|
+
If this metadata is missing, the result is not reproducible enough to count.
|
|
62
|
+
|
|
63
|
+
`rendered` runs must list at least one repository-local screenshot file for both baseline and AI Design Context outputs. Each screenshot must be a regular image file stored inside its matching `baseline/` or `ai-design-rules/` directory. Preview links can be supplementary, but do not replace stored visual evidence. `directional` runs remain useful for learning but must record their limitation.
|
|
64
|
+
|
|
65
|
+
## Scoring Model
|
|
66
|
+
|
|
67
|
+
Use `docs/EVALUATION_RUBRIC.md`.
|
|
68
|
+
|
|
69
|
+
Every dimension is scored from `0` to `10`.
|
|
70
|
+
|
|
71
|
+
Total score:
|
|
72
|
+
|
|
73
|
+
```text
|
|
74
|
+
sum(all dimensions) / number of dimensions
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
Report:
|
|
78
|
+
|
|
79
|
+
- baseline total;
|
|
80
|
+
- AI Design Context total;
|
|
81
|
+
- absolute difference;
|
|
82
|
+
- per-category differences;
|
|
83
|
+
- reviewer notes.
|
|
84
|
+
|
|
85
|
+
Do not collapse scores into a single claim without showing category scores.
|
|
86
|
+
|
|
87
|
+
## Reproducibility Requirements
|
|
88
|
+
|
|
89
|
+
To make results comparable:
|
|
90
|
+
|
|
91
|
+
- use the exact scenario file without editing the brief;
|
|
92
|
+
- use the same LLM for both runs;
|
|
93
|
+
- use the same implementation target for both runs;
|
|
94
|
+
- use the same time budget for both runs;
|
|
95
|
+
- use the same tool access for both runs;
|
|
96
|
+
- use the same evaluator or at least two independent evaluators;
|
|
97
|
+
- publish raw outputs and scores;
|
|
98
|
+
- do not fix one output manually before scoring.
|
|
99
|
+
|
|
100
|
+
If implementation fails, record the failure and score the observable output.
|
|
101
|
+
|
|
102
|
+
## Limitations
|
|
103
|
+
|
|
104
|
+
This benchmark cannot remove all subjectivity.
|
|
105
|
+
|
|
106
|
+
Known limitations:
|
|
107
|
+
|
|
108
|
+
- design scoring has reviewer judgment;
|
|
109
|
+
- different models may respond differently to long context;
|
|
110
|
+
- implementation quality can affect perceived design quality;
|
|
111
|
+
- small scenarios do not prove enterprise-scale validity;
|
|
112
|
+
- current rules and patterns are seed-level, not complete.
|
|
113
|
+
|
|
114
|
+
Benchmark results should guide rule, pattern, prompt, and skill evolution. They should not be used as marketing claims unless the evidence is public and reproducible.
|
|
115
|
+
|
|
116
|
+
## Result Storage
|
|
117
|
+
|
|
118
|
+
Store future results in `evidence/`.
|
|
119
|
+
|
|
120
|
+
Do not add synthetic scores. Do not invent studies. Do not compare against private internal systems.
|
|
121
|
+
|
|
122
|
+
## Next Rendered Run: Todo
|
|
123
|
+
|
|
124
|
+
Use `benchmarks/todo-app.md` for the first paired rendered run of the next release. It exercises the updated Daily Home Surface (`PAT-001`, `VIS-002`) and Context-Preserving Preview (`PAT-005`, `UX-004`) alongside existing capture, accessibility, and recovery guidance. This is a run protocol, not a completed experiment or evidence of improvement.
|
|
125
|
+
|
|
126
|
+
The [first execution on 2026-09-28](../evidence/todo/2026-09-28-codex-paired-todo/README.md) stores a completed pair and an internal unblinded evaluation. Keep this protocol for repeats; the recorded result does not establish general improvement.
|
|
127
|
+
|
|
128
|
+
### Freeze The Setup
|
|
129
|
+
|
|
130
|
+
Before either generation starts, record one shared setup in the run notes:
|
|
131
|
+
|
|
132
|
+
- The scenario file, its hash, the knowledge revision, and any uncommitted knowledge patch used by the rules run.
|
|
133
|
+
- The exact common brief: the Product Brief, Expected User, Expected Platform, Constraints, and Core User Tasks sections from `benchmarks/todo-app.md`, unchanged and identical in both prompts.
|
|
134
|
+
- One model/provider/version, generation surface, exposed sampling settings, tool set, and time budget. Record unexposed settings as unexposed; do not invent values.
|
|
135
|
+
- The implementation target: static HTML, CSS, and JavaScript served locally, with no build dependencies or external assets. Use the same target for both runs.
|
|
136
|
+
- The same deterministic task data, browser/version, viewport sizes, and storage-reset procedure for both evaluations.
|
|
137
|
+
|
|
138
|
+
Give both runs the same technical requirement to document how the evaluator can reproduce empty, loading, failed-save, and successful-save states. Local test hooks are acceptable; do not expose test controls as product features. Let each run design its interface from the same brief.
|
|
139
|
+
|
|
140
|
+
Use two fresh generation sessions with no shared conversation or output access. Give the baseline only the common brief and shared technical setup. Give the rules session those same inputs plus the pinned knowledge source and instructions to retrieve context for `daily-home-surface`, `quick-capture`, and `context-preserving-preview`, with `--platform mobile`, before implementation. Save the exact returned context and actual prompts. Both sessions must have equivalent tool access; document any isolation limitations.
|
|
141
|
+
|
|
142
|
+
Do not use the existing runnable Todo fixture as either generated result. Do not manually improve one output before scoring. A failure to build, render, or reproduce a required state is a recorded result.
|
|
143
|
+
|
|
144
|
+
### Capture The Same Tasks
|
|
145
|
+
|
|
146
|
+
Evaluate both outputs at **390 × 844** and **1440 × 900** CSS pixels, at the same browser zoom and device scale. Reset storage and use the same task data before each sequence. Save these local artifacts inside each run type's `screenshots/` directory:
|
|
147
|
+
|
|
148
|
+
| Artifact suffix | Action and observable evidence |
|
|
149
|
+
| --- | --- |
|
|
150
|
+
| `default.png` | Open the populated daily surface; inspect current work and the primary action. |
|
|
151
|
+
| `empty.png` | Reset to no tasks; inspect the starting action. |
|
|
152
|
+
| `saving.png` | Submit a task under controlled delay; inspect input retention and layout stability. |
|
|
153
|
+
| `save-error.png` | Cause a failed save; inspect error text, retained input, and retry. |
|
|
154
|
+
| `saved.png` | Retry successfully; inspect feedback and duplicate prevention. |
|
|
155
|
+
| `completed.png` | Mark a task complete; inspect state clarity and remaining work. |
|
|
156
|
+
| `detail.png` | Open task details after scrolling; inspect the relation to the source list. |
|
|
157
|
+
| `keyboard-focus.png` | Reach a core action by keyboard; inspect visible focus and accessible naming. |
|
|
158
|
+
| `detail-reduced-motion.png` | Repeat detail opening with reduced motion; inspect equivalent information and controls. |
|
|
159
|
+
|
|
160
|
+
Prefix each filename with `mobile-` or `desktop-`. Record the exact reproduction steps with each capture. If a state cannot be reached, record it as missing; do not create a substitute screenshot that pretends to show the state.
|
|
161
|
+
|
|
162
|
+
Screenshots do not prove timing, focus restoration, or interruption behavior. Record interaction results separately: close the preview and verify source position/focus; close or change selection during opening and check for stale content; repeat with reduced motion; verify retry and task completion. Performance observations remain qualitative unless measurements are actually collected.
|
|
163
|
+
|
|
164
|
+
### Score And Publish Evidence
|
|
165
|
+
|
|
166
|
+
Use the same evaluator and `docs/EVALUATION_RUBRIC.md` for both outputs. Score visible behavior, include a note and artifact reference for each category, and report all category differences as well as the averages. If possible, score anonymous A/B outputs before revealing their conditions. Disclose whether evaluation was blinded and whether the evaluator participated in generation; a parent agent reviewing its own team's output is not independent evaluation.
|
|
167
|
+
|
|
168
|
+
Store completed runs using `evidence/TEMPLATE.md`, including raw source, prompts, settings, screenshots, state coverage, evaluator identity, and limitations. Run `npm run benchmark:validate`. The validator currently checks artifact presence, metadata consistency, and image signatures; it does not enforce this state matrix or judge visual quality. The evaluator must check those manually.
|
|
169
|
+
|
|
170
|
+
Do not create scored evidence entries or change rule maturity while preparation is incomplete. One completed pair remains a small-sample result; it cannot establish universal improvement or isolate the effect of context retrieval from pattern changes. Repeated pairs or a separate previous-version comparison are needed for those stronger claims.
|
|
@@ -0,0 +1,110 @@
|
|
|
1
|
+
# DesignLint Readiness
|
|
2
|
+
|
|
3
|
+
DesignLint v0 is implemented as a small, deterministic relationship checker.
|
|
4
|
+
|
|
5
|
+
This document defines its current scope and the richer validation that can follow metadata migration.
|
|
6
|
+
|
|
7
|
+
## Inputs
|
|
8
|
+
|
|
9
|
+
DesignLint should consume:
|
|
10
|
+
|
|
11
|
+
- `schema/*.schema.json` for object validation;
|
|
12
|
+
- `registry/objects.json` for ID resolution;
|
|
13
|
+
- `registry/relationships.json` for graph edges;
|
|
14
|
+
- YAML front matter from knowledge objects after migration;
|
|
15
|
+
- generated indexes after registry-backed indexing exists.
|
|
16
|
+
|
|
17
|
+
## Required Metadata
|
|
18
|
+
|
|
19
|
+
DesignLint requires stable metadata:
|
|
20
|
+
|
|
21
|
+
- `id`
|
|
22
|
+
- `alias`
|
|
23
|
+
- `slug`
|
|
24
|
+
- `title`
|
|
25
|
+
- `object_type`
|
|
26
|
+
- `status`
|
|
27
|
+
- `version`
|
|
28
|
+
- `category`
|
|
29
|
+
- `relationships`
|
|
30
|
+
- `last_reviewed_at`
|
|
31
|
+
- `maturity`
|
|
32
|
+
- `risk_level`
|
|
33
|
+
|
|
34
|
+
Optional metadata such as `platform`, `product_type`, `surface`, `applies_to`, and `does_not_apply_to` will improve context matching.
|
|
35
|
+
|
|
36
|
+
## DesignLint v0
|
|
37
|
+
|
|
38
|
+
Run it with:
|
|
39
|
+
|
|
40
|
+
```bash
|
|
41
|
+
npm run lint:design
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
It reads only registered objects and typed registry relationships. It fails when:
|
|
45
|
+
|
|
46
|
+
- a rule has no upstream research via `derived_from` or `inspired_by`;
|
|
47
|
+
- a pattern does not `require` a rule;
|
|
48
|
+
- a prompt has no pattern, rule, or checklist reference;
|
|
49
|
+
- a checklist has no rule or pattern coverage;
|
|
50
|
+
- a reference project does not implement a prompt and validate a prompt or pattern;
|
|
51
|
+
- a review does not require a checklist and validate a knowledge object.
|
|
52
|
+
|
|
53
|
+
It does not judge prose, visual quality, or whether a relationship is substantively well chosen. `npm run validate` remains responsible for synchronizing registry relationships with front matter.
|
|
54
|
+
|
|
55
|
+
## Future Validation
|
|
56
|
+
|
|
57
|
+
With the schema-first foundation, DesignLint can later validate:
|
|
58
|
+
|
|
59
|
+
- object metadata matches the correct schema;
|
|
60
|
+
- IDs are globally unique;
|
|
61
|
+
- aliases are not reused within the same object type;
|
|
62
|
+
- slugs are file-safe;
|
|
63
|
+
- relationship types are allowed;
|
|
64
|
+
- relationship targets exist;
|
|
65
|
+
- deprecated objects are not used by active prompts;
|
|
66
|
+
- patterns require rules;
|
|
67
|
+
- prompts require patterns and rules;
|
|
68
|
+
- reference projects validate patterns and prompts;
|
|
69
|
+
- rules trace back to research;
|
|
70
|
+
- orphan objects are detected;
|
|
71
|
+
- manual indexes match registry output.
|
|
72
|
+
|
|
73
|
+
## Current Validation Layer
|
|
74
|
+
|
|
75
|
+
`tools/validate-knowledge.mjs` is the first lightweight validation step.
|
|
76
|
+
|
|
77
|
+
It currently validates:
|
|
78
|
+
|
|
79
|
+
- YAML front matter on migrated research, rules, patterns, prompts, checklists, reviews, and reference projects;
|
|
80
|
+
- required metadata fields;
|
|
81
|
+
- object ID prefixes for migrated object types;
|
|
82
|
+
- unique aliases and slugs within the migrated set;
|
|
83
|
+
- registry coverage for migrated files;
|
|
84
|
+
- registry relationship source and target IDs;
|
|
85
|
+
- front matter relationship targets that point to registered IDs or existing files;
|
|
86
|
+
- duplicate IDs, aliases, slugs, and raw JSON keys;
|
|
87
|
+
- local Codex skill metadata and generated-index freshness.
|
|
88
|
+
|
|
89
|
+
This is a metadata migration guard. It complements DesignLint v0 rather than replacing it.
|
|
90
|
+
|
|
91
|
+
## Not Implemented Yet
|
|
92
|
+
|
|
93
|
+
The current repository does not yet include:
|
|
94
|
+
|
|
95
|
+
- registry-backed observations;
|
|
96
|
+
- schema validation in CI;
|
|
97
|
+
- graph traversal;
|
|
98
|
+
- stale-object detection;
|
|
99
|
+
- relationship-quality evaluation;
|
|
100
|
+
- visual or runtime product evaluation.
|
|
101
|
+
|
|
102
|
+
## Readiness Standard
|
|
103
|
+
|
|
104
|
+
Before DesignLint rules are implemented, the repository should have:
|
|
105
|
+
|
|
106
|
+
- metadata front matter on every knowledge object;
|
|
107
|
+
- complete registry coverage;
|
|
108
|
+
- generated indexes;
|
|
109
|
+
- relationship validation;
|
|
110
|
+
- a stable ID allocation process.
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
# Design System
|
|
2
|
+
|
|
3
|
+
AI Design Context is not a design system or component library.
|
|
4
|
+
|
|
5
|
+
This document defines how agents should think when a product already has design-system structure.
|
|
6
|
+
|
|
7
|
+
## Order Of Decisions
|
|
8
|
+
|
|
9
|
+
```text
|
|
10
|
+
tokens -> themes -> components -> screens
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
## Agent Guidance
|
|
14
|
+
|
|
15
|
+
- Reuse existing tokens before creating new visual values.
|
|
16
|
+
- Use semantic names for color, spacing, typography, and state.
|
|
17
|
+
- Do not create one-off styles when an existing token or component fits.
|
|
18
|
+
- Do not treat components as the product model.
|
|
19
|
+
- Let rules and patterns decide behavior; let the design system keep execution consistent.
|
|
20
|
+
|
|
21
|
+
## Boundary
|
|
22
|
+
|
|
23
|
+
AI Design Context can describe how design-system decisions should be used by agents. It does not ship production components, token packages, or UI kit assets.
|